Method and System for Ensuring Data Consistency

By linking multiple KV blocks in the KV chain, the problem of data writing twice in the prior art is solved, and the effect of reducing the number of writes, reducing overhead and IO congestion is achieved.

CN112540994BActive Publication Date: 2025-06-20SAMSUNG ELECTRONICS CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010985710.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-04-08
Filing Date
2020-09-18
Publication Date
2025-06-20
Estimated Expiration
2040-09-18

AI Technical Summary

Technical Problem

In order to ensure data consistency, existing databases and file systems need to write data twice, resulting in additional overhead and input and output (IO) congestion and adding write amplification factor (WAF).

Method used

By linking multiple KV blocks in a key value (KV) chain to ensure data consistency, the specific method includes assigning the internal key to the first KV block and restoring the start internal key, assigning the next internal key different from the internal key, and encapsulating the corresponding user key value in the first KV block and the next KV block.

Benefits of technology

Reduces the number of writes required to ensure data consistency, reduces system overhead and IO congestion, reduces write amplification factor, and improves data storage efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112540994B_ABST
    Figure CN112540994B_ABST
Patent Text Reader

Abstract

A method and system for ensuring data consistency are provided. The method includes: assigning an internal key to both a first KV block and a recovery start internal key, assigning a next internal key different from the internal key and corresponding to a next KV block, and encapsulating corresponding user key values in the first KV block and the next KV block, wherein the first KV block is accessed by reading the recovery start internal key, and wherein the next KV block is accessed by reading the next internal key of the first KV block.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims priority and the benefit of U.S. Provisional Application No. 62 / 903,651, filed on September 20, 2019, entitled "Reliable Key-Value Store With Write-Ahead-Log-Less Mechanism", the entire content of which is incorporated herein by reference. Technical Field

[0002] One or more embodiments of the present disclosure generally relate to data storage. Background Art

[0003] Some databases and file systems write the same data twice to ensure consistency. For example, some key-value (KV) stores use a write-ahead log (WAL) as a mechanism for ensuring data consistency to support the reliability of the KV store.

[0004] That is, in some databases, before updating the actual data and metadata for helping data recovery in the case of system crash / power outage (system failure), the WAL is initially used to write the KV block, which may be a significant bottleneck for supporting system reliability. Thereafter, when updating the metadata corresponding to the KV block, the KV block is written a second time.

[0005] For consistency, related databases and file systems also write data twice. For example, related database management systems can use double writes with a storage engine, while other systems can use a journaling file system.

[0006] Therefore, by writing data twice, the system experiences additional overhead (e.g., doubling the bandwidth used), which causes input / output (IO) congestion and increases the write amplification factor (WAF). Summary of the Invention

[0007] The embodiments described herein provide an improvement to data storage.

[0008] According to some embodiments of the present disclosure, a method for chaining a plurality of KV blocks in a KV chain to ensure data consistency is provided. The method includes: assigning an internal key to both a first KV block and a recovery start internal key, assigning a next internal key different from the internal key and corresponding to a next KV block, and encapsulating corresponding user key values in the first KV block and the next KV block, wherein the first KV block is accessed by reading the recovery start internal key, and wherein the next KV block is accessed by reading the next internal key of the first KV block.

[0009] The method may further include: updating all user key values of a first KV block, generating or updating a metadata table to reference the first KV block, marking the first KV block as eligible for deletion from the KV device, and updating the recovery start internal key to the next internal key of the first KV block.

[0010] The method may further include: statically allocating a recovery start key, wherein the recovery start internal key is accessed by reading the recovery start key.

[0011] The method may further include, after a system failure of a KV device storing the first KV block and the next KV block: reading a recovery start key from the KV device, obtaining a recovery start internal key using the recovery start key, locating and reading the first KV block using the recovery start internal key, obtaining the next internal key of the next KV block from the first KV block, and locating and reading the next KV block using the next internal key.

[0012] The method may further include: reading an additional next internal key that is part of the device value of the next KV block, the additional next internal key corresponding to a subsequent next KV block, locating and reading the subsequent next KV block using the additional next internal key, and repeating until the corresponding next internal key corresponds to a KV block not found in the KV device.

[0013] The method may further include: determining that the first KV block does not have valid user key values, marking the first KV block as eligible for deletion, and updating the recovery start internal key to the next internal key of the first KV block, the next internal key corresponding to a subsequent KV block.

[0014] The next internal key may include a part of the device value in the first KV block and include a device key for the next KV block.

[0015] According to other embodiments of the present disclosure, a system for ensuring data consistency by chaining a plurality of KV blocks in a KV chain is provided, the system including a key-value storage engine configured to: assign an internal key to both a first KV block and a recovery start internal key, assign a next internal key that is different from the internal key and corresponds to a next KV block, and encapsulate corresponding user key values in the first KV block and the next KV block, wherein the first KV block is accessed by reading the recovery start internal key, and wherein the next KV block is accessed by reading the next internal key of the first KV block.

[0016] The key-value storage engine can also be configured to: update all user key-values of the first KV block, generate or update a metadata table to reference the first KV block, mark the first KV block as eligible for deletion from the KV device, and update the recovery start internal key to the next internal key of the first KV block.

[0017] The key-value storage engine can also be configured to statically allocate a recovery start key, wherein the recovery start internal key is accessed by reading the recovery start key.

[0018] The key-value storage engine can also be configured to, after a system failure of the KV device storing the first KV block and the next KV block: read the recovery start key from the KV device, obtain the recovery start internal key using the recovery start key, use the recovery start internal key to locate and read the first KV block, obtain the next internal key of the next KV block from the first KV block, and use the next internal key to locate and read the next KV block.

[0019] The key-value storage engine can also be configured to: read an additional next internal key that is part of the device value of the next KV block, the additional next internal key corresponding to a subsequent next KV block, use the additional next internal key to locate and read the subsequent next KV block, and repeat until the corresponding next internal key corresponds to a KV block not found in the KV device.

[0020] The key-value storage engine can also be configured to: determine that the first KV block does not have a valid user key, mark the first KV block as eligible for deletion, and update the recovery start internal key to the next internal key of the first KV block, the next internal key corresponding to a subsequent KV block.

[0021] The next internal key can include a part of the device value in the first KV block and include a device key for the next KV block.

[0022] According to yet another other embodiment of the present disclosure, there is provided a non-transitory computer-readable medium implemented on a system for linking a plurality of KV blocks in a KV chain to ensure data consistency, the non-transitory computer-readable medium having computer code that, when executed on a processor, implements a method of data storage, the method including: allocating an internal key to both the first KV block and a recovery start internal key, allocating a next internal key that is different from the internal key and corresponds to a next KV block, and encapsulating corresponding user key-values in the first KV block and the next KV block, wherein the first KV block is accessed by reading the recovery start internal key, and wherein the next KV block is accessed by reading the next internal key of the first KV block.

[0023] When executed on a processor, the computer code can also implement a method of data storage by the following operations: updating all user key values of a first KV block, generating or updating a metadata table to reference the first KV block, marking the first KV block as eligible for deletion from the KV device, and updating the recovery start internal key to the next internal key of the first KV block.

[0024] When executed on a processor, the computer code can also implement a method of data storage by statically allocating a recovery start key, wherein the recovery start internal key is accessed by reading the recovery start key.

[0025] When executed on a processor after a system failure of a KV device storing the first KV block and the next KV block, the computer code can also implement a method of data storage by the following operations: reading the recovery start key from the KV device, using the recovery start key to obtain the recovery start internal key, using the recovery start internal key to locate and read the first KV block, obtaining the next internal key of the next KV block from the first KV block, and using the next internal key to locate and read the next KV block.

[0026] When executed on a processor, the computer code can also implement a method of data storage by the following operations: reading an additional next internal key that is part of the device value of the next KV block, the additional next internal key corresponding to a subsequent next KV block, using the additional next internal key to locate and read the subsequent next KV block, and repeating until the corresponding next internal key corresponds to a KV block not found in the KV device.

[0027] When executed on a processor, the computer code can also implement a method of data storage by the following operations: determining that the first KV block does not have a valid user key, marking the first KV block as eligible for deletion, and updating the recovery start internal key to the next internal key of the first KV block, the next internal key corresponding to a subsequent KV block.

[0028] Accordingly, the system of embodiments of the present disclosure can improve data storage by reducing the number of writes required to ensure data consistency. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Figure 1 is a block diagram depicting the overall concept of a WAL-free recovery mechanism according to an embodiment of the present disclosure;

[0030] Figure 2A is a block diagram depicting a mechanism for ensuring system consistency;

[0031] Figure 2B is a block diagram depicting a WAL-free recovery mechanism for ensuring system consistency according to an embodiment of the present disclosure;

[0032] Figure 3 is a block diagram depicting a method of generating a KV chain in a KV device according to an embodiment of the present disclosure;

[0033] Figure 4 is a block diagram depicting a method of generating a KV block to be used in a KV chain in a KV device; and

[0034] Figure 5 is a block diagram depicting a method of updating a metadata table and starting an iKey in a KV device according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0035] The features of the inventive concept and the method of implementing the inventive concept can be more easily understood through a detailed description of the embodiments with reference to the accompanying drawings. Hereinafter, the embodiments will be described in more detail with reference to the drawings. However, the described embodiments can be implemented in various different forms and should not be construed as being limited to the embodiments shown herein. Instead, these embodiments are provided as examples so that the present disclosure will be thorough and complete, and will fully convey the aspects and features of the inventive concept to those skilled in the art. Accordingly, processes, elements, and techniques that are not necessary for a person of ordinary skill in the art to fully understand the aspects and features of the inventive concept may not be described.

[0036] Unless otherwise noted, the same reference numerals represent the same elements throughout the drawings and the written description, and thus, their description will not be repeated. In addition, parts that are not relevant to the description of the embodiments may not be shown to make the description clear. In the drawings, the relative dimensions of elements, layers, and regions may be exaggerated for clarity.

[0037] In the detailed description, for purposes of explanation, numerous specific details are set forth to provide a thorough understanding of the various embodiments. However, it is apparent that the various embodiments can be practiced without these specific details or with one or more equivalent arrangements. In other instances, well-known structures and devices are shown in block diagram form to avoid unnecessarily obscuring the various embodiments.

[0038] It will be understood that although the terms “first,” “second,” “third,” etc. may be used herein to describe various elements, components, regions, layers, and / or portions, these elements, components, regions, layers, and / or portions should not be limited by these terms. These terms are used to distinguish one element, component, region, layer, or portion from another element, component, region, layer, or portion. Thus, a first element, component, region, layer, or portion described below may be referred to as a second element, component, region, layer, or portion without departing from the spirit and scope of the present disclosure.

[0039] The terms used herein are for the purpose of describing particular embodiments only and are not intended to limit the present disclosure. As used herein, unless the context clearly indicates otherwise, the singular forms "a" and "an" are also intended to include the plural forms. It will also be understood that when the terms "comprises," "comprising," "has," and "having" and variations thereof are used in this specification, they specify the presence of the stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items.

[0040] As used herein, the terms "substantially," "about," "approximate," and similar terms are used as terms of approximation and not of degree, and are intended to account for the inherent deviations of measured or calculated values recognized by a person of ordinary skill in the art. As used herein, "about" or "approximate" includes the stated value and means within an acceptable deviation range of the particular value determined by a person of ordinary skill in the art in view of the measurements discussed and the errors associated with the measurement of a particular quantity (i.e., the limitations of the measurement system). For example, "about" can mean within one or more standard deviations, or within ±30%, 20%, 10%, 5% of the stated value. Additionally, the use of "may" in describing embodiments of the present disclosure refers to "one or more embodiments of the present disclosure."

[0041] When embodiments can be implemented differently, the specific processing sequences can be performed in an order different from the described order. For example, two consecutively described processes can be performed substantially simultaneously or in an order opposite to the described order.

[0042] An electronic or electrical device and / or any other related device or component according to embodiments of the present disclosure described herein can be implemented using any suitable hardware, firmware (e.g., an application specific integrated circuit), software, or a combination of software, firmware, and hardware. By way of example, the various components of these devices can be formed on one integrated circuit (IC) chip or on different IC chips. Additionally, the various components of these devices can be implemented on a flexible printed circuit film, a tape carrier package (TCP), a printed circuit board (PCB), or formed on a substrate.

[0043] In addition, various components of these devices can be processes or threads running on one or more processors in one or more computing devices, which execute computer program instructions and interact with other system components for performing the various functions described herein. The computer program instructions are stored in a memory, which can be implemented in a computing device using standard memory devices (such as, by way of example, random access memory (RAM)). The computer program instructions can also be stored in other non-transitory computer-readable media (such as, by way of example, CD-ROMs, flash drives, etc.). In addition, those skilled in the art should recognize that, without departing from the spirit and scope of the embodiments of the present disclosure, the functions of various computing devices can be combined or integrated into a single computing device, or the functions of a particular computing device can be distributed across one or more other computing devices.

[0044] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this inventive concept belongs. It will also be understood that terms, such as those defined in commonly used dictionaries, should be interpreted as having a meaning that is consistent with their meaning in the relevant art and / or the context of this specification, and should not be interpreted in an idealized or overly formal sense.

[0045] Embodiments of the present disclosure enable a storage device to construct recovery information without using a write-ahead log (WAL), thereby improving the field of data storage by reducing overhead, input / output (IO) congestion, and / or write amplification factor (WAF).

[0046] Figure 1 is a block diagram depicting the general concept of a WAL-less recovery mechanism according to an embodiment of the present disclosure.

[0047] Referring to Figure 1 , a recovery mechanism that avoids the need for a write-ahead log (WAL) according to an embodiment of the present disclosure can generally operate as described herein. A "start key-value (KV) block (e.g., a start KV block for recovery)" 160 can be stored in a database / file system (e.g., Figure 2A and Figure 2B the database / file system 210 shown in Figure 2B ) and written to a corresponding KV device (e.g., Figure 3The KV storage engine 350) links the start KV block 160 to the first KV block 180a by including the internal key (iKey) of the first KV block 180a (e.g., the written KV block (1)) in the start KV block 160. That is, the KV storage engine maintains the start KV block 160 and updates the start KV block 160 whenever the first KV block 180a changes. Similarly, the KV storage engine links the first KV block 180a to the second KV block 180b by including the iKey of the second KV block 180b (e.g., the written KV block (2)) in the first KV block 180a. The second KV block 180b is then further linked to the third KV block 180c (e.g., the written KV block (3)) in a similar manner, and the third KV block 180c in turn is linked to the fourth KV block 180d (e.g., the written KV block (4)), and so on. Thus, the first KV block 180a to the fourth KV block 180d form a KV chain and are included in the recovery range 140 in the event of a system failure. It should be noted that even when a block is updated, any of the KV blocks 180a to 180d within the recovery range 140 cannot be deleted. Instead, to prevent truncating the KV chain, only keys outside the recovery range 140 can be deleted. In addition, the recovery range 140 can be expanded later using upcoming KV blocks.

[0048] Thereafter, for example, the key values of the first KV block 180a and the second KV block 180b can be updated or written to one or more corresponding KV devices by the database / file system. Then, the metadata table 170 can be updated to indicate that the first KV block 180a and the second KV block 180b have been successfully updated or written (e.g., the metadata table {(1), (2)}). Once the metadata table 170 is updated, the start KV block 160 can be updated to include the iKey of the third KV block 180c to alternatively link to the third KV block 180c. Accordingly, the recovery range 140 can be updated to exclude the first KV block 180a and the second KV block 180b. Then the first KV block 180a and the second KV block 180b can be marked as meeting the deletion conditions.

[0049] Thus, in the event of a system failure, as recorded in the metadata table 170, since the first KV block 180a and the second KV block 180b have been successfully updated or written before the system failure, the recovery can start by updating the start KV block 160 to correspond to the third KV block 180c without redundantly updating the first KV block 180a and the second KV block 180b, such that the recovery range 140 starts at the third KV block 180c.

[0050] Figure 2A is a block diagram depicting a mechanism for ensuring system consistency.

[0051] Refer toFigure 2A , in the example system 200a, the database / file system 210 and the corresponding block device 220a are initialized, and one or more files / keys 230 (e.g., file / key A, file / key B) can be submitted to the database / file system 210 in a write request 240 so that the corresponding data 250 (e.g., user key values) (e.g., data A, data B) are written to the block device 220a.

[0052] Thereafter, the database / file system 210 writes to the write-ahead log (WAL) / journal file 260 in the block device 220a so that the data 250 and the metadata 270 corresponding to the data 250 (e.g., metadata {A}, metadata {B}, metadata {A, B}) are written to the WAL / journal file 260.

[0053] Then, the database / file system 210 updates or writes the data 250 and the metadata 270 a second time to the storage device of the block device 220a (e.g., an area in the block device 220a different from the WAL / journal file 260). Finally, after the data 250 and the metadata 270 are successfully written to the block device 220a, the data 250 and the metadata 270 can be deleted from the WAL / journal file 260.

[0054] Figure 2B is a block diagram depicting a WAL-less recovery mechanism for ensuring system consistency according to an embodiment of the present disclosure.

[0055] Referring to Figure 2B , in the system 200b according to an embodiment of the present disclosure, similar to the system 200a using the above block device 220a, the database / file system 210 and the KV device 220b are initialized, and one or more files / keys 230 can be submitted to the database / file system 210 for a write request 240 so that the corresponding data 250 (e.g., user key values) are written to the KV device 220b.

[0056] However, different from Figure 2A the system 200a that uses a write-ahead log (WAL) / journal file implemented based on the block device 220a, the system 200b of various embodiments alternatively implements based on the KV device. Figure 2B The system 200b of

[0057] In addition, different from the above-mentioned system 200a, the database / file system 210 of the system 200b according to the embodiments of the present disclosure encapsulates data 250 (e.g., data A, data B) and files / keys 230 (e.g., file / key A, file / key B), as well as iKeys (e.g., A, B, C) included in one or more corresponding KV blocks 280 for indicating subsequent KV blocks (further described below with reference to Figure 3 ). Further, the system 200b links the KV blocks 280 together in a KV chain, and then writes the KV blocks 280 to the KV device 220b by using the iKeys indicating the next KV block. Each KV block 280 can be connected or linked together in the KV chain by including the iKey of the subsequent KV block 280 within the data of the immediately preceding KV block 280 (e.g., as part of the device value). The database / file system 210 can also link the recovery start block 290 to the first KV block by including the iKey of the first KV block in the KV block 280 in the recovery start block 290, so as to enable the start of recovery after a system failure.

[0058] Once it is determined that the encapsulated data 250 and the iKeys of the subsequent KV blocks 280 have been successfully written to the KV device 220b, the database / file system 210 can update the metadata 270 (e.g., metadata {A, B}) corresponding to the written encapsulated data 250 in the KV device 220b to indicate that the KV blocks 280 have been successfully written and to indicate the positions of the KV blocks 280. Then, the database / file system 210 can update the recovery start block 290 by using the iKeys of the subsequent KV blocks 280, and the metadata 270 corresponding to the subsequent KV blocks 280 has not been updated in the KV device 220b.

[0059] For example, even if there is another KV block 280 downstream in the KV chain that has been successfully written to the KV device 220b and its corresponding metadata 270 has been updated, the first KV block 280 in the KV chain that has not been successfully written to the KV device 220b or its corresponding metadata 270 has not been updated can also correspond to the iKey updated in the recovery start block 290, thereby ensuring data consistency in the case of a system failure.

[0060] Therefore, as shown above, according to the disclosed embodiments, the WAL / log file 260 of the system 200a can be omitted.

[0061] Figure 3 is a block diagram depicting a method of generating a KV chain in a KV device according to an embodiment of the present disclosure.

[0062] Referring to Figure 3, user 310 may seek to write one or more user key values 330 to a KV store / database (e.g., in KV device 320). According to the user's write request, KV storage engine 350 inserts user key value 330 into KV block 380 as part of the device value of the first KV block 380a. That is, user key value 330 can be encapsulated by the corresponding KV block 380.

[0063] KV storage engine 350 may also allocate iKey 340a as the device key for the first KV block 380a. However, it should be noted that in some embodiments, iKey 340 can be a user key rather than a device key. Additionally, iKey 340 can be a number or a string, as long as iKey 340 is unique (even for the same user key).

[0064] The iKey 340a of the first KV block 380a can be recorded as the start iKey 360 in KV device 320 (e.g., the recovery start internal key) to enable retrieving recovery data. At the same time, the recovery start key 365 can be used as the key for the recovery data. The recovery start key 365 can be statically allocated by KV storage engine 350 and can remain unchanged. In contrast, as will be described below, the start iKey 360 is the value of the recovery data including the recovery start key 365 and can be updated.

[0065] In addition to user key value 330, KV storage engine 350 may also insert the value of iKey 340b for the second KV block 380b into the first KV block 380a as another part of the device value of the first KV block 380a. That is, each KV block 380 may also include a "next iKey" 345 for indicating the subsequent KV block 380, and the subsequent KV block 380 may not have been created when the next iKey 345 is inserted into the KV block 380. The next iKey 345, like iKey 340, can be pre-allocated when the database is established.

[0066] The above process can be repeated for multiple KV blocks (e.g., where the iKey 340c for the third KV block 380c is included as part of the device value, and this part serves as the next iKey 345b for the second KV block 380b, etc.). Thus, each KV block 380 typically contains a unique iKey 340 that can be assigned on the host side, one or more user key values 330, and a next iKey 345 for indicating the next KV block 380, where the next iKey 345a of the first KV block 380a will reference the iKey 340b of the second KV block 380b in a manner similar to how the start iKey 360 references the iKey 340a of the first KV block 380a in the KV chain. Therefore, as will be further discussed below, the order of the KV blocks 380 in the KV chain is embedded in the KV blocks 380, and the next iKey 345 enables recovery after a system failure.

[0067] By having each KV block 380 include the above information, each KV block 380 has a link for linking itself to the next KV block 380 (e.g., the subsequent KV block), thereby forming a KV chain (e.g., a chain of KV blocks 380). For the last KV block 380 in the KV chain, even if the iKey 340 for the next KV block 380 does not exist in the KV device 320, an iKey is still assigned as the iKey 340 for that next KV block. Thus, if a KV block 380 does not exist in the KV device 320, the end point of recovery has been reached.

[0068] It should be noted that since the KV blocks 380 do not need to be physically contiguous blocks, the next iKey 345 does not necessarily reference a physically adjacent KV block 380. For example, in some embodiments, the next KV block 380 can be determined by chronological order rather than "key" order. For example, the key order can be 1->2->3->4->5, while the write order (or chronological order) can be 3->1->2->5->4. Thus, the KV chain order will be 3->1->2->5->4.

[0069] Thus, one or more KV blocks 380 can be recovered and can be written to the KV device 320 after a system failure by using the recovery start key 365 corresponding to each KV block 380 and the respective pre-allocated iKey 340. Thus, the KV device 320 ensures atomic writes using single KV block atomicity.

[0070] Figure 4 is a block diagram depicting a method of generating KV blocks to be used in a KV chain in a KV device according to an embodiment of the present disclosure.

[0071] In summary, with reference to Figure 4, when establishing the database, the initially pre-allocated iKey 442 is used as the iKey 440 of the current KV block 480. In addition, the iKey 440 of the first KV block 480 in the KV chain is stored as the start iKey in the KV device (e.g., Figure 3 the start iKey 360 in the KV device 320 of

[0072] ). Then, since the initially pre-allocated iKey 442 has been used for the current KV block 480 and a new iKey will be used for the subsequent / next KV block 480, the initially pre-allocated iKey 442 cannot be used for the subsequent KV block. Therefore, the iKey generator 490 can be used to generate a new iKey to update the pre-allocated iKey 442. The current KV block 480 can use the initially pre-allocated iKey 442.

[0073] Thus, the pre-allocated iKey 442 can be used to generate a new pre-allocated iKey 442a such that the updated pre-allocated iKey 442a can be used as the iKey for the subsequently created KV block. For example, the iKey generator 490 can generate a new pre-allocated iKey 442a, and then the pre-allocated iKey 442a can be updated. The new pre-allocated iKey 442a can also be stored in the "internal key for the next KV block" (e.g., the next iKey 445). That is, the iKey generator 490 generates the pre-allocated iKey 442 that serves as the iKey 440 for the KV block 480, and then generates another pre-allocated iKey 442a as the next iKey 445 corresponding to the iKey of the next (future) KV block.

[0074] That is, the iKey generator 490 can generate a unique new iKey that is not subordinate to the previously mentioned initially pre-allocated iKey 442. In addition, the unique new iKey can be stored in both the updated pre-allocated iKey 442a and the next iKey 445 for the next KV block, such that the updated pre-allocated iKey 442a and the next iKey 445 are the same iKey.

[0075] The iKey generator 490 can generate only a single iKey at a time. Then, the generated iKey can be used / updated / written to both the pre-allocated iKey 442a and the next iKey 445. Additionally, the initially used iKey 440 includes the previously allocated iKey 440 linked to the initially pre-allocated iKey 442. However, the embodiment is not limited thereto, and the internal iKey generator 490 can generate multiple iKeys for one or more KV chains at a time.

[0076] Then, the iKey generator 490 can add the updated pre-allocated iKey 442a to the KV block 480 as the next iKey 445 corresponding to the subsequent KV block, and also add the updated iKey 442a as the subsequent iKey for the subsequent KV block. That is, the updated pre-allocated iKey 442a will be used for both the iKey of the next KV block and the next iKey 445 of the current KV block 480.

[0077] Thereafter, after generating and inserting the initial iKey 440 for the current block 480 and the next iKey 445 for the next KV block 480, the user key value 430 can be inserted into the KV block 480. However, the embodiment is not limited thereto, and the initial iKey 440, the next iKey 445, and the user key value 430 can be inserted into the KV block 480 simultaneously or in an order different from the order described in the present disclosure.

[0078] Figure 5 is a block diagram depicting a method of updating a metadata table and a start iKey (e.g., a resume start internal key) in a KV device according to an embodiment of the present disclosure. In Figure 5 where n is an integer greater than 100.

[0079] Referring to Figure 5 , the operation of constructing the KV block 580 can be shown. For example, a brief overview of constructing the KV block 580 is as follows. First, an iKey can be appended, then, an iKey for the next KV block (e.g., the next iKey 545) can be appended, and thereafter, a user key and value (e.g., the user KV 530) can be appended to complete the KV block 580. Each KV block 580 includes a next iKey 545 of the iKey 540 linked to the subsequent KV block 580, and the KV blocks 580 (e.g., the KV blocks 580 corresponding to reference numerals 100, 101, 102, 104, 105,..., n) are linked to form a KV chain.

[0080] If the metadata 570 of one or more KV blocks 580 (e.g., up to the metadata 570 of the KV block 580 corresponding to reference numeral 99) is stored in the corresponding metadata table 575, the starting iKey 560 will be changed to the iKey 540 of the earliest KV block 580 in the KV chain that is not indicated in the metadata table 575 (e.g., the KV block 580 corresponding to reference numeral 100). Since the KV blocks 580 referenced in the metadata table 575 can be assumed to have been successfully written to the KV device 520 and can be recovered without referring to the KV chain in the event of a system failure, the starting iKey 560 can be updated. In contrast, if the KV block 580 is not referenced in the metadata table 575, the KV block 580 cannot be found without referring to the starting iKey 560.

[0081] After the starting iKey 560 is updated, the KV blocks 580 indicated in the metadata table 575 and before the KV block 580 corresponding to the updated starting iKey 560 can be marked for deletion. Additionally, the KV blocks 580 after the KV block 580 corresponding to the starting iKey 560 or the KV blocks 580 not referenced in the metadata table 575 cannot be deleted (e.g., even if the user deletes all the keys in the KV block 580 (e.g., all the user keys in the KV block 580 are deleted and become invalid user keys) or updates all the keys in the KV block 580), so as to ensure that the intact KV chain can be maintained as a linked list.

[0082] For example, in Figure 5 , the KV blocks 580a, 580b, and 580c can be deleted after the user deletes all the corresponding key values 530 in these KV blocks 580a, 580b, and 580c (e.g., all the corresponding user key values 530 in these KV blocks 580 are deleted and become invalid user keys) or updates all the corresponding key values 530 in these KV blocks 580a, 580b, and 580c. However, since the intermediate KV blocks 580d and 580e are not referenced in the metadata table 575, even if the KV block 580f is indicated as being included in the metadata table 575 and is indicated as having been successfully written to the KV device 520, the KV block 580f cannot be deleted. This is to ensure that the KV chain of the remaining KV blocks 580 is not damaged. That is, since the starting iKey 560 is updated to point to the KV block 580d before the KV block 580f in the KV chain, the KV block 580f remains undeleted to ensure that the KV chain remains undamaged and to ensure that all the KV blocks 580 in the metadata table that have not been updated or referenced can be recovered in the event of a system failure.

[0083] In the case of a system failure, the access recovery start key 565 is accessed, and the start iKey 560 is obtained from the KV device 520 to initiate recovery after the system failure. Then, by using the start iKey 560, the corresponding KV block 580 that is the first in the KV chain is read. By reading the first KV block 580 of the KV chain, the iKey 540 of the next KV block 580 can be obtained. Thus, the next KV block 580 can be found and read. This process can be repeated until the KV block 580 of the next iKey 545 read does not exist in the KV device 520. After all KV blocks 580 that are not written to the metadata table 575 are obtained, the metadata table 575 can be updated, and the start iKey 560 can be changed. However, the embodiment is not limited thereto. After a part of the KV blocks 580 among all KV blocks 580 that are not written to the metadata table 575 are obtained, the metadata table 575 can be updated, and the start iKey 560 can be changed. For example, after one or more KV blocks 580 among all KV blocks 580 that are not written to the metadata table 575 are obtained, the metadata table 575 can be updated, and the start iKey 560 can be changed.

[0084] Because the KV blocks 580 are linked by the links of the KV chain, all KV blocks 580 that are not written to the metadata table 575 can be located by simply accessing the recovery start key 565 to use the start iKey 560 to locate the first KV block 580 of the remaining KV chain. Therefore, the second redundant write can be avoided without any risk to data consistency. Additionally, contrary to the list of all unwritten KV blocks 580 (e.g., contrary to the WAL / log file 260 shown in Figure 2A ), only the start iKey 560 is stored.

[0085] Therefore, the embodiments of the present disclosure provide a KV chain with single or multiple user key values to avoid double writes for KV storage consistency, and only keys outside the recovery range can be deleted to prevent truncating the KV chain. If the system has multiple write / commit threads / workers, multiple recovery start keys 565 can be saved in the device. Each write / commit thread should have at least one recovery start key 565 in order to have its own write stream. Otherwise, the chain may be broken because the device cannot guarantee the write order.

[0086] Note that the KV device guarantees the atomicity for writing only one KV block. In addition, the embodiments of the present disclosure enable recovery of the KV storage (or system) from a system failure by using the KV chain mechanism. Furthermore, the embodiments of the present disclosure can be implemented by a system that assigns device keys (iKeys) to KV blocks and forms a KV chain using the KV blocks for recovery, wherein the user keys are embedded in the values of the corresponding KV blocks.

[0087] Advantages provided by the disclosed embodiments include no double writing of data and no WAL for crash recovery. Accordingly, the increased WAF and increased IO bandwidth due to double writing can be reduced, the effective throughput can be increased, and recovery can be supported without sacrificing effective IO performance.

Claims

1. A method for ensuring data consistency, the method comprising: Assign an internal key to both the first key-value block and the resume start internal key; Assign a next internal key, which is different from the internal key and corresponds to the next key-value block, to the first key-value block and the next key-value block; And Encapsulate the corresponding user key values in the first key-value block and the next key-value block, Wherein, the first key-value block is accessed by reading the resume start internal key, Wherein, the next key-value block is accessed by reading the next internal key of the first key-value block.

2. The method according to claim 1, the method further comprising: Update all user key values of the first key-value block; Generate or update a metadata table to reference the first key-value block; Mark the first key-value block as meeting the condition for deletion from the key-value device; And Update the resume start internal key to the next internal key of the first key-value block.

3. The method according to claim 1, the method further comprising: Static allocate a resume start key, wherein the resume start internal key is accessed by reading the resume start key.

4. The method according to claim 3, the method further comprising, after a system failure of a key - value device storing a first key - value block and the next key - value block: Reading a recovery start key from the key - value device; Use the resume start key to obtain the resume start internal key; Use the resume start internal key to locate and read the first key-value block; Obtain the next internal key of the next key-value block from the first key-value block; And Use the next internal key to locate and read the next key-value block.

5. The method according to claim 4, the method further comprising: Locate and read the subsequent next key-value block, wherein the step of locating and reading the subsequent next key-value block includes: reading an additional next internal key, which is part of the device value of the next key-value block and corresponds to the subsequent next key-value block, and using the additional next internal key to locate and read the subsequent next key-value block; and Repeat the step of locating and reading the subsequent next key-value block until the corresponding next internal key corresponds to a key-value block not found in the key-value device.

6. The method according to any one of claims 1 to 5, the method further comprising: Determine that the first key-value block does not have valid user keys; Mark the first key-value block as meeting the deletion condition; And Update the resume start internal key to the next internal key of the first key-value block, which corresponds to the subsequent key-value block.

7. The method according to any one of claims 1 to 5, wherein, The next internal key is part of the device value in the first key-value block and is the device key for the next key-value block.

8. A system for ensuring data consistency, the system comprising: A key-value storage engine, configured to: Assign an internal key to both the first key-value block and the resume start internal key; Assign a next internal key, which is different from the internal key and corresponds to the next key-value block, to the first key-value block and the next key-value block; And Encapsulate the corresponding user key values in the first key-value block and the next key-value block, Wherein, the first key-value block is accessed by reading the resume start internal key, Wherein, the next key-value block is accessed by reading the next internal key of the first key-value block.

9. The system according to claim 8, wherein, The key-value storage engine is further configured to: Update all user key values of the first key-value block; Generate or update a metadata table to reference the first key-value block; Mark the first key-value block as meeting the condition for deletion from the key-value device; And Update the resume start internal key to the next internal key of the first key-value block.

10. The system according to claim 8, wherein, The key-value storage engine is further configured to: statically allocate a resume start key, wherein the resume start internal key is accessed by reading the resume start key.

11. The system according to claim 10, wherein, The key-value storage engine is further configured to, after a system failure of the key-value device storing the first key-value block and the next key-value block: Read the resume start key from the key-value device; Obtain a recovery start internal key using a recovery start key; Locate and read a first key value block using the recovery start internal key; Obtain the next internal key of the next key value block from the first key value block; And Locate and read the next key value block using the next internal key.

12. The system according to claim 11, wherein, The key value storage engine is further configured to: Locate and read subsequent next key value blocks, wherein the process of locating and reading subsequent next key value blocks includes: reading an additional next internal key that is part of the device value of the next key value block, the additional next internal key corresponding to the subsequent next key value block, and using the additional next internal key to locate and read the subsequent next key value block; and Repeat the process of locating and reading subsequent next key value blocks until the corresponding next internal key corresponds to a key value block not found in the key value device.

13. The system according to any one of claims 8 to 12, wherein, The key value storage engine is further configured to: Determine that the first key value block does not have a valid user key; Mark the first key value block as meeting the deletion condition; And Update the recovery start internal key to the next internal key of the first key value block, the next internal key corresponding to a subsequent key value block.

14. The system according to any one of claims 8 to 12, wherein, The next internal key is part of the device value in the first key value block and is the device key for the next key value block.

15. A non - transitory computer - readable medium implemented on a system for ensuring data consistency, the non - transitory computer - readable medium having computer instructions that, when executed on a processor, implement a method for data storage, the method comprising: Assign an internal key to both the first key value block and the recovery start internal key; Assign a next internal key that is different from the internal key and corresponds to the next key value block to the first key value block and the next key value block; And Encapsulate the corresponding user key values in the first key value block and the next key value block, wherein the first key value block is accessed by reading the recovery start internal key, wherein the next key value block is accessed by reading the next internal key of the first key value block.

16. The non-transitory computer-readable medium according to claim 15, wherein, The computer instructions, when executed on a processor, further implement a method of data storage by: Updating all user key values of the first key value block; Generating or updating a metadata table to reference the first key value block; Marking the first key value block as meeting the condition for deletion from the key value device; And Updating the recovery start internal key to the next internal key of the first key value block.

17. The non-transitory computer-readable medium according to claim 15, wherein, The computer instructions, when executed on a processor, further implement a method of data storage by statically assigning a recovery start key, wherein the recovery start internal key is accessed by reading the recovery start key.

18. The non-transitory computer-readable medium according to claim 17, wherein, After a system failure of a key value device storing the first key value block and the next key value block, when the computer instructions are executed on a processor, a method of data storage is further implemented by: Reading a recovery start key from the key value device; Obtaining a recovery start internal key using the recovery start key; Locating and reading the first key value block using the recovery start internal key; Obtaining the next internal key of the next key value block from the first key value block; And Locating and reading the next key value block using the next internal key.

19. The non-transitory computer-readable medium according to claim 18, wherein, The computer instructions, when executed on a processor, further implement a method of data storage by: Locate and read the next key-value block that follows, wherein the steps of locating and reading the next key-value block that follows include: reading an additional next internal key that is part of the device value of the next key-value block, the additional next internal key corresponding to the next key-value block that follows, and using the additional next internal key to locate and read the next key-value block that follows; and Repeat the steps of locating and reading the next key-value block that follows until the corresponding next internal key corresponds to a key-value block not found in the key-value device.

20. The non-transitory computer-readable medium according to any one of claims 15 to 19, wherein, The computer instructions, when executed on a processor, also implement a data storage method by the following operations: Determine that the first key-value block does not have a valid user key; Mark the first key-value block as meeting the deletion criteria; And Update the restore start internal key to the next internal key of the first key-value block, the next internal key corresponding to the subsequent key-value block.

Citation Information

Patent Citations

  • Method and device for storing data

    CN106886375A

  • A memory system and a data storage method

    CN109918352A