Data storage method and system for performing a merge with transaction grouping

By identifying and grouping transactions and merging conflicting data writing in transaction groups, unnecessary updates and data inconsistency caused by multi-transaction updates in key-value stores are solved, and more efficient data storage and stronger data consistency are achieved.

CN112540859BActive Publication Date: 2025-05-13SAMSUNG ELECTRONICS CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010987021.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-03-30
Filing Date
2020-09-18
Publication Date
2025-05-13
Estimated Expiration
2040-09-18

AI Technical Summary

Technical Problem

In key-value storage operations, multiple transactions may update the same key value at the same time, resulting in unnecessary key-value updates and data inconsistency problems. Especially when the system crashes, a large number of transactions need to be rolled back to ensure data consistency.

Method used

By identifying multiple transactions in the pending queue and their corresponding key value updates, identifying common association keys for common association key value updates, grouping transactions based on the transaction group ID, and combining conflicting data in the transaction group, updating metadata only after all merged key value updates are written to the storage device.

Benefits of technology

Reduces the number of transactions that need to be rolled back after a system failure, improves the writing efficiency and data consistency of data storage, avoids unnecessary overwriting and merging, and enhances the stability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112540859B_ABST
    Figure CN112540859B_ABST
Patent Text Reader

Abstract

A data storage method and a system for performing a merge with transaction grouping are provided, the method comprising: identifying a plurality of transactions in a suspended queue, the plurality of transactions having one or more key value updates respectively corresponding to a plurality of keys; identifying a common associated key among the plurality of keys that is associated with a common associated key value update among key value updates belonging to different transactions among the plurality of transactions; assigning transaction group IDs to the plurality of transactions respectively based on respective transaction IDs assigned to the transaction group IDs; grouping the plurality of transactions into corresponding transaction groups among a plurality of transaction groups based on the assigned transaction group IDs; and merging conflicting data writes corresponding to the common associated key value update of the common associated key for grouped transactions among the plurality of transactions in the same transaction group among the plurality of transactions.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims priority to and the benefit of U.S. Provisional Application No. 62 / 903,662, filed on September 20, 2019, entitled “Transaction Grouping for Overwrite Merging,” the contents of which are incorporated herein in their entirety. Technical Field

[0002] One or more aspects of embodiments of the present disclosure relate generally to data storage. Background Art

[0003] In the operation of a key-value store, a given key may be written multiple times, and each key may be updated substantially simultaneously. Multiple transactions for updating each key value may be placed in a suspended queue. Different transactions within the suspended queue may correspond to conflicting key value updates, all of which correspond to a common key. For example, a first write operation or first transaction may attempt to perform a key value update on a key, and thereafter a second write operation / second transaction attempts to perform a different key value update on the same key.

[0004] As each transaction is processed, each key value update is processed in turn. Thus, a particular key may be updated with a key value, and only updated again with a different key value shortly thereafter. That is, although a later transaction appears to be unrelated to a key value update of an earlier transaction before the initial key value update occurs, both key value updates may occur, one being unnecessary. In such cases, different transactions typically occur relatively close to each other in time. Summary of the invention

[0005] Embodiments described herein provide improvements to data storage.

[0006] According to an embodiment of the present disclosure, a method for data storage is provided, the method comprising: identifying multiple transactions in a suspended queue, the multiple transactions having one or more key value updates respectively corresponding to multiple keys; identifying a common associated key among the multiple keys that is associated with a common associated key value update among key value updates belonging to different transactions among the multiple transactions; assigning transaction group IDs to the multiple transactions respectively based on corresponding transaction IDs assigned to the transaction group IDs; grouping the multiple transactions into corresponding transaction groups among multiple transaction groups based on the assigned transaction group IDs; and merging conflicting data writes corresponding to the common associated key value update of the common associated key for grouped transactions among the multiple transactions in the same transaction group among the multiple transactions.

[0007] When a crash occurs at a crashing transaction among transactions in a crashing transaction group among the plurality of transaction groups, the method may further include retrying only linked transactions among transactions in the crashing transaction group that are linked to the crashing transaction due to merging of conflicting data writes.

[0008] The method may also include removing transactions from the multiple transactions that only have one or more earlier key value updates, wherein the one or more earlier key value updates appear to be unrelated to one or more later key value updates corresponding to the same respective key, wherein the later key value updates are in one or more other transactions from the multiple transactions that occur later in time than the removed transaction from the multiple transactions.

[0009] The method may also include determining a pair of consecutive transactions lacking any commonly associated key value updates corresponding to a common key among the plurality of keys; and assigning transaction group IDs to the consecutive transactions, respectively, so that the consecutive transactions are located in different corresponding transaction groups.

[0010] The method may also include: determining which pair of consecutive transactions has the least number of commonly associated key value updates corresponding to a common key among the plurality of keys; and assigning transaction group IDs to the consecutive transactions, respectively, so that the consecutive transactions are located in different corresponding transaction groups.

[0011] Assigning a transaction group ID may also be based on an analysis of the total number of commonly associated key value updates across different transactions.

[0012] The method may also include: writing all merged key value updates written to conflicting data of a transaction group among the multiple transaction groups to one or more storage devices; and updating metadata corresponding to the merged key value updates of the transaction group among the multiple transaction groups only when it is confirmed that all merged key value updates have been written to the one or more storage devices.

[0013] The method may further include assigning a transaction ID to the plurality of transactions.

[0014] According to another embodiment of the present disclosure, a system for performing a merge with transaction grouping is provided, the system comprising: a transaction module for: identifying a plurality of transactions in a suspended queue, the plurality of transactions having one or more key value updates respectively corresponding to a plurality of keys, identifying a common associated key among the plurality of keys that is associated with a common associated key value update among key value updates belonging to different transactions among the plurality of transactions, assigning transaction group IDs to the plurality of transactions respectively based on corresponding transaction IDs assigned to the transaction group IDs, and grouping the plurality of transactions into corresponding transaction groups among a plurality of transaction groups based on the assigned transaction group IDs; and a merge module for merging conflicting data writes corresponding to the common associated key value update of the common associated key for grouped transactions among the plurality of transactions in the same transaction group among the plurality of transactions.

[0015] When a crash occurs at a crashed transaction among transactions in a crashed transaction group among the plurality of transaction groups, the transaction module may be further configured to retry only linked transactions among transactions in the crashed transaction group that are linked to the crashed transaction due to merging of conflicting data writes.

[0016] The merge module may also be configured to remove transactions from the multiple transactions that have only one or more earlier key-value updates, wherein the one or more earlier key-value updates appear to be unrelated to one or more later key-value updates corresponding to the same respective key, wherein the later key-value updates are in one or more other transactions from the multiple transactions that occur later in time than the removed transaction from the multiple transactions.

[0017] The transaction module may also be configured to: determine a pair of consecutive transactions lacking any commonly associated key value updates corresponding to a common key among the plurality of keys; and assign transaction group IDs to the consecutive transactions, respectively, so that the consecutive transactions are located in different corresponding transaction groups.

[0018] The transaction module may also be configured to: determine which pair of consecutive transactions has the least number of commonly associated key value updates corresponding to a common key among the plurality of keys; and assign transaction group IDs to the consecutive transactions respectively so that the consecutive transactions are located in different corresponding transaction groups.

[0019] The transaction module may be further configured to assign the transaction group ID based on an analysis of a total number of commonly associated key value updates across different transactions.

[0020] The system may also include: an in-flight request buffer configured to: write all merged key value updates of conflicting data written by a transaction group among the multiple transaction groups to one or more storage devices; and update metadata corresponding to the merged key value updates of the transaction group among the multiple transaction groups only when it is confirmed that all merged key value updates have been written to the one or more storage devices.

[0021] The transaction module may be further configured to assign transaction IDs to the plurality of transactions.

[0022] According to another embodiment of the present disclosure, a non-transitory computer-readable medium implemented on a system for performing a merge with transaction grouping is provided, the non-transitory computer-readable medium having a computer code, the computer code implementing a method of data storage when executed on a processor, the method comprising: identifying a plurality of transactions in a suspended queue, the plurality of transactions having one or more key value updates respectively corresponding to a plurality of keys; identifying a common association key among the plurality of keys that is associated with a common association key value update among key value updates belonging to different transactions among the plurality of transactions; assigning transaction group IDs to the plurality of transactions respectively based on corresponding transaction IDs assigned to the transaction group IDs; grouping the plurality of transactions into corresponding transaction groups among a plurality of transaction groups based on the assigned transaction group IDs; and merging conflicting data writes corresponding to the common association key value update of the common association key for grouped transactions among the plurality of transactions in the same transaction group among the plurality of transactions.

[0023] When a crash occurs at a crashed transaction in transactions in a crashed transaction group in the multiple transaction groups, the computer code, when executed on a processor, can also implement the method for data storage by the following steps: only retrying transactions in the crashed transaction group that are linked to the crashed transaction due to merging of conflicting data writes.

[0024] When the computer code is executed on a processor, the method for data storage can also be implemented by the following steps: removing transactions in the multiple transactions that only have one or more earlier key value updates in time, wherein the one or more earlier key value updates in time appear to be unrelated to one or more later key value updates in time corresponding to the same corresponding key, and the later key value updates in time are in one or more other transactions in the multiple transactions that occur later in time than the removed transaction in the multiple transactions.

[0025] When the computer code is executed on a processor, the method for data storage can also be implemented by the following steps: determining a pair of consecutive transactions that lack any common associated key value updates corresponding to a common key among the multiple keys; and assigning transaction group IDs to the consecutive transactions respectively, so that the consecutive transactions are located in different corresponding transaction groups.

[0026] Therefore, the system of the embodiments of the present disclosure can improve data storage by reducing the number of transactions that must be rolled back (retried by the system) to ensure data consistency after a system failure. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Non-limiting and non-exhaustive embodiments of the present invention are described with reference to the following figures, wherein like reference numerals refer to like parts throughout the various figures unless otherwise specified.

[0028] Figure 1 is a block diagram depicting an overwrite merge and the transaction dependencies resulting therefrom;

[0029] Figure 2 is a block diagram depicting overwrite merging with transaction grouping according to some embodiments of the present disclosure;

[0030] Figure 3 is a flowchart depicting a method of overwrite merging with transaction grouping according to an embodiment of the present disclosure; and

[0031] Figure 4 is a block diagram depicting a workflow according to an embodiment of the present disclosure.

[0032] Throughout the several figures of the accompanying drawings, corresponding reference numerals indicate corresponding elements. It will be appreciated by those skilled in the art that the elements in the accompanying drawings are shown for simplicity and clarity and are not necessarily drawn to scale. For example, the dimensions of some of the elements, layers, and regions in the accompanying drawings may be exaggerated relative to other elements, layers, and regions to help improve clarity and understanding of the various embodiments. In addition, common but well-understood elements and parts that are not relevant to the description of the embodiments may not be shown to facilitate less obstructed viewing of these various embodiments and to make the description clear. DETAILED DESCRIPTION

[0033] By referring to the specific implementation and the accompanying drawings of the embodiments, the features of the inventive concept and the methods for realizing the inventive concept can be more easily understood. Hereinafter, the embodiments will be described in more detail with reference to the accompanying drawings. However, the described embodiments can be implemented in various different forms and should not be construed as being limited to the embodiments shown herein. On the contrary, these embodiments are provided as examples so that the present disclosure will be thorough and complete, and the aspects and features of the inventive concept will be fully conveyed to those skilled in the art. Therefore, unnecessary processing, elements and techniques for a person of ordinary skill in the art to fully understand the aspects and features of the inventive concept may not be described.

[0034] Unless otherwise specified, the same reference numerals denote the same elements throughout the drawings and written description, and therefore their description will not be repeated. In addition, parts not related to the description of the embodiments may not be shown to make the description clear. In the drawings, the relative sizes of elements, layers, and regions may be exaggerated for clarity.

[0035] In the detailed description, for the purpose of explanation, many specific details are set forth to provide a thorough understanding of the various embodiments. However, it is clear that the various embodiments can be practiced without these specific details or with one or more equivalent arrangements. In other instances, well-known structures and devices are shown in block diagram form to avoid unnecessarily obscuring the various embodiments.

[0036] It will be understood that, although the terms "first", "second", "third", etc. may be used herein to describe various elements, components, regions, layers and / or parts, these elements, components, regions, layers and / or parts should not be limited by these terms. These terms are only used to distinguish one element, component, region, layer or part from another element, component, region, layer or part. Therefore, without departing from the spirit and scope of the present disclosure, the first element, first component, first region, first layer or first part discussed below may be referred to as the second element, second component, second region, second layer or second part.

[0037] The terms used herein are only used to describe the purpose of specific embodiments, and are not intended to limit the present disclosure. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. It will also be understood that when the terms "comprise", "have" and "include" are used in this specification, the features, integral bodies, steps, operations, elements and / or components stated are explained, but the presence or addition of one or more other features, integral bodies, steps, operations, elements, components and / or their groups are not excluded. As used herein, the term "and / or" includes any combination and all combinations of one or more of the associated listed items.

[0038] As used herein, the terms "substantially," "about," "approximately," and similar terms are used as approximate terms rather than terms of degree, and are intended to take into account the inherent deviations in measured or calculated values ​​that one of ordinary skill in the art will recognize. As used herein, "about" or "approximately" includes the stated value and means within an acceptable range of deviations of the particular value determined by one of ordinary skill in the art taking into account the measurement in question and the errors associated with the measurement of the particular quantity (i.e., the limitations of the measurement system). For example, "about" may mean within one or more standard deviations, or within ±30%, 20%, 10%, 5% of the stated value. In addition, "may" is used when describing embodiments of the present disclosure to mean "one or more embodiments of the present disclosure."

[0039] When a specific embodiment can be implemented differently, a specific processing order can be performed differently from the order described. For example, two consecutively described processes can be performed substantially simultaneously, or in the reverse order of the described order.

[0040] Any suitable hardware, firmware (e.g., an application specific integrated circuit), software, or a combination of software, firmware, and hardware may be used to implement the electronic or electrical devices and / or any other related devices or components according to the embodiments of the present disclosure described herein. For example, the various components of these devices may be formed on an integrated circuit (IC) chip or on separate IC chips. In addition, the various components of these devices may be implemented on a flexible printed circuit film, a tape carrier package (TCP), a printed circuit board (PCB), or formed on a substrate.

[0041] In addition, the various components of these devices can be processes or threads that run on one or more processors in one or more computing devices, execute computer program instructions and interact with other system components for performing the various functions described herein. Computer program instructions are stored in a memory that can be implemented in a computing device using a standard memory device (e.g., a random access memory (RAM)). Computer program instructions may also be stored in other non-temporary computer-readable media (e.g., a CD-ROM, a flash drive, etc.). In addition, those skilled in the art should recognize that without departing from the spirit and scope of the embodiments of the present disclosure, the functions of various computing devices may be combined or integrated into a single computing device, or the functions of a particular computing device may be distributed in one or more other computing devices.

[0042] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as those commonly understood by those skilled in the art to which the inventive concept belongs. It will also be understood that, unless explicitly defined as such herein, terms (such as those defined in general dictionaries) should be interpreted as having a meaning consistent with their meaning in the context of this specification and / or the relevant art, and should not be interpreted in an idealized or overly formal sense.

[0043] In order to improve write efficiency in a storage device by reducing the number of key value updates, "inline overwrite merging" or "overwrite merging", which may be referred to as "in-place updating", enables merging of conflicting writes that correspond to a common key and occur within a given time frame. The conflicting writes correspond to different corresponding transactions that occur within the given time frame, such that the two transactions are substantially present in the pending queue at the same time.

[0044] Thus, overwrite merging allows older key value updates to a particular key to be discarded in favor of newer key value updates to the same key. That is, conflicting pending writes to a common key that occur within a given time frame can be "merged" by discarding unrelated earlier writes in favor of the latest writes in time. Thus, the total number of key value updates / writes can be reduced, thereby reducing system overhead.

[0045] The concept of overwrite merging can be applied as a write reduction method associated with various KV storage methods (e.g., a key-value store used in conjunction with a key-value solid-state drive (KVSSD) that is used in conjunction with an existing persistent key-value store (such as a Developed by , and is compatible with RocksDB For example, one of the bases of the concept of overwrite merging is the ability to write multiple threads sequentially, where each thread creates keys that duplicate the keys of other threads. Therefore, when flash memory is used in the storage media space, adding the ability for simple overwrite merging may be beneficial, and the reduction of write amplification may be beneficial.

[0046] Although overwrite merging can be used to reduce write amplification, overwrite merging is not usually used alone for transactional systems. For example, when used in combination with transactional batch processing, overwrite merging / in-place updates may have unwanted effects. Such unwanted effects may be caused by the long tail of suspended transactions caused by transaction dependencies, thereby potentially causing data inconsistency and making overwrite merging impractical. In addition, data consistency may be potentially destroyed. That is, although overwrite merging can reduce write amplification and improve performance (for example, improve input and output operations per second (IOPS)), overwrite merging may not be feasible when there are a large number of transactions. For example, when there are a large number of transactions, when a crash occurs during one of a plurality of cascaded / linked transactions, data consistency may be destroyed. In addition, when multiple overwrites occur, the long tail of suspended transactions may occur.

[0047] Thus, embodiments of the present disclosure enable: grouping cascaded transactions (e.g., using transaction groups) by allowing inline overwrite merging. To implement the disclosed embodiments, writing of corresponding metadata may be delayed until all transaction operations in the group are written to the device. As noted, although inline overwrite merging according to the presented embodiments may be used to reduce write amplification, inline overwrite merging may not be used alone for transactional systems. By utilizing flash memory technology to perform embodiments of the present disclosure, some flash memory devices may exhibit improved write endurance and performance (e.g., IO operations per second (IOPS)).

[0048] Figure 1 is a block diagram depicting an override merge and the transaction dependencies that result from it.

[0049] Reference Figure 1 , the first transaction Trxn1 may have a key value update V2, which may be overwritten and merged with the key value update V3 of the second transaction Trxn2. The second transaction Trxn2 may have a key value update V4, which may be overwritten and merged with the key value update V5 of the third transaction Trxn3 and the key value update V6 of the fourth transaction Trxn4, respectively. The third transaction Trxn3 may also have a key value update V7, which may be overwritten and merged with the key value update V8 of the fourth transaction Trxn4. In addition, the fourth transaction Trxn4 may have a key value update V9, which may be overwritten and merged with the key value update V10 of the fifth transaction Trxn5.

[0050] Therefore, due to the interdependence of the transactions (due to the ability of their key value updates to be overwritten and merged), multiple transactions can be linked together through the resulting chain of interdependencies. As a result, in the event of a crash or unexpected power outage, undesirable effects caused by overwrite merging may occur.

[0051] In this example, if Figure 1 As shown in , if a crash occurs after the key value update V9 of the fourth transaction Trxn4 and before the key value update V10 of the fifth transaction Trxn5, all linked transactions Trxn1 to Trxn5 may be rolled back (i.e., reattempted by the system to ensure validity and data consistency). That is, due to the crash, the key value update V9 of the fourth transaction Trxn4 cannot be determined to be valid. Because the key value update V9 of the fourth transaction Trxn4 cannot be determined to be valid, the completion of the fourth transaction Trxn4 cannot be determined to be valid.

[0052] In addition, because the key value update V6 of the fourth transaction Trxn4 is overwritten and merged with the key value update V4 of the second transaction Trxn2 and the key value update V5 of the third transaction Trxn3, the key value updates V4 and V5 cannot be confirmed as having been written effectively. Therefore, the second transaction Trxn2 and the third transaction Trxn3 cannot be determined to be effective similarly.

[0053] Finally, since the key value update V2 of the first transaction Trxn1 and the key value update V3 of the second transaction Trxn2 are overwritten and merged, and since the second transaction Trxn2 cannot be determined to be valid, the first transaction Trxn1 cannot be determined to be valid and must also be rolled back.

[0054] Therefore, due to transaction dependencies caused by overwrite merging, a crash may require the system to retry a relatively large number of transactions in order to achieve data consistency.

[0055] Furthermore, there may be situations where transactions are unrecoverable after a crash. For example, there may be situations where the metadata for key-value update V1 has already been written to storage, and there is no record of the previous metadata in any key-value store implementation (e.g., see the . Figure 3 S380 and Figure 4 Table 460 in ), so the first transaction Trxn1 cannot be rolled back.

[0056] Embodiments of the present disclosure can avoid problems that may potentially arise from overwrite merging by combining overwrite merging with a concept referred to herein as "transaction grouping," that is, by grouping transactions via transaction grouping, transactional writes to the same key can be reliably merged even in the event of a crash, thereby improving write efficiency.

[0057] Figure 2 is a block diagram depicting overwrite merging with transaction grouping according to some embodiments of the present disclosure.

[0058] For example, refer to Figure 2 , different transactions may be grouped by assigning corresponding transaction group identifiers (IDs) (e.g., Group1 or Group2) to multiple transactions. The transaction group ID may be generated from one or more corresponding transaction IDs. In addition, if no overwrite merge occurs, a new transaction group ID may be started. For example, if the number of updates exceeds a threshold, a new transaction group ID may be started.

[0059] By generating a corresponding transaction group ID for each transaction based on the transaction ID of each transaction, one or more transaction group demarcations that divide each transaction group may be determined. As an example, in one implementation of this embodiment, the transaction group ID may be the most significant 59 bits of the transaction ID. If a new transaction group ID is to be used, the next transaction ID may be incremented.

[0060] exist Figure 2 In the example shown in , consecutive transactions are included in a single transaction group, but the disclosed embodiments are not limited thereto. In addition, the specific number of transactions in each transaction group is not particularly limited.

[0061] exist Figure 2 In the example shown in , the first to third transactions Trxn1, Trxn2, and Trxn3 are each assigned a first transaction group ID to be placed in the first group "Group1", and the fourth transaction Trxn4 and the fifth transaction Trxn5 are assigned a second transaction group ID to be placed in the second group "Group2". Therefore, although the third transaction Trxn3 and the fourth transaction Trxn4 include corresponding key value updates that both correspond to the same corresponding key (e.g., key value updates V5 and V6 corresponding to key "C", and key value updates V7 and V8 corresponding to key "D"), since the third transaction Trxn3 and the fourth transaction Trxn4 belong to different transaction groups Group1 and Group2, respectively, these key value updates will not be overwritten and merged.

[0062] Furthermore, as described above, grouping can address situations where transactions after a crash are conversely unrecoverable. That is, conventionally, when the first transaction Trxn1 cannot be rolled back because the metadata of the key value update V1 has already been written to the storage device and there is no record of the previous metadata in any key value storage implementation, the transaction may be unrecoverable. However, grouping according to an embodiment of the present disclosure avoids such an unrecoverable crash because the metadata of the updated key value can only be updated when the metadata of the updated key value is written to the storage device to ensure data consistency in the event of a crash (e.g., see the description further below). Figure 3 S380 and Figure 4 Table 460 in ).

[0063] Therefore, similar to the above Figure 1 Unlike the example described, if the crash occurs after the key value update V9 of the fourth transaction Trxn4 and before the key value update V10 of the fifth transaction Trxn5, only the linked transactions Trxn4 and Trxn5, which are part of the same transaction group Group2, are rolled back. Since the transactions Trxn1, Trxn2, and Trxn3 of the first transaction group Group1 are confirmed to be valid, they do not need to be rolled back as a result of the crash.

[0064] That is, by dividing the consecutive transactions Trxn3 and Trxn4 into different transaction groups, the link between the consecutive transactions Trxn3 and Trxn4 that would otherwise occur due to overwrite merging does not exist / is removed. Therefore, the third transaction Trxn3 can be determined to be valid, and the inefficiency caused by the crash can be reduced.

[0065] In general, the advantages of transaction grouping with overwrite coalescing provided by embodiments disclosed herein include improved write efficiency in the event of a crash. As described above, if a crash occurs, the system may not need to roll back beyond the beginning of a given transaction group associated with the crash.

[0066] Furthermore, an advantage of transaction grouping with override merging provided by embodiments disclosed herein may be achieved by preventing override merging between different adjacent transactions belonging to different corresponding transaction groups while still allowing override merging across different transactions that are commonly assigned a given transaction group ID to be within a corresponding transaction group.

[0067] Thus, the system may not have to roll back, as it would have if there were no transaction grouping.However, the system is still able to achieve the advantages associated with overwrite merging, at least to some extent, by overwriting some of the merging key value updates.

[0068] exist Figure 2 In the example shown in , it can be noted that having to write to keys C and D twice each time (key value updates V5 and V6 for key C, and key value updates V7 and V8 for key D) can cause inefficiencies. Therefore, in some embodiments, transaction group boundaries can be assigned by evaluating which consecutive transactions (if any) do not have any key value updates that can be overwritten and merged with each other. In other embodiments, transaction group boundaries can be assigned by evaluating which consecutive transactions have the fewest key value updates that are eligible to be overwritten and merged with each other.

[0069] Furthermore, in other embodiments of the present disclosure, the system may allow for adjustment of the generation of transaction group IDs based on analysis of common corresponding keys corresponding to key value updates across different transactions.

[0070] For example, in Figure 2In the example shown in , none of the key updates from the third transaction Trxn3 are used (V5 and V7 of Trxn3 are both rendered unrelated to V6 and V8 of Trxn4). Therefore, the system can evict Trxn3 from the first transaction group Group1. For example, the key object has a link to the key updates, and each key update object contains a transaction ID. Background threads or "write workers" can be used to check the key updates and find possible overwrite merges. These background threads can create new transaction group IDs (for example, if the number of updates exceeds a threshold).

[0071] Figure 3 is a flow chart depicting a method of overwrite merging with transaction grouping according to an embodiment of the present disclosure.

[0072] Reference Figure 3 According to an embodiment of the present disclosure, the disclosed system may initially determine whether an incoming transaction may be too large to successfully perform an overwrite merge with transaction grouping (S310). If yes, a transaction ID and a subsequent transaction group ID may be assigned to the transaction (S320). If no, a transaction ID and a current transaction group ID may be assigned to the transaction (S330).

[0073] After the transaction ID and the transaction group ID are assigned to the transaction, a transaction operation count may be added to the group (S340). That is, the transaction operations divided into the same transaction group may be counted, and / or the transaction operations that have been completed in the same transaction group may be counted to determine whether the transaction group is full and / or whether all transaction operations in the transaction group have been completed in subsequent processing. Then, an inline overwrite merge may be performed on the transactions including the incoming transaction (S350). Thereafter, a write operation may be performed to write the data corresponding to the transaction to one or more devices (S360).

[0074] The disclosed system may then determine whether all operations in the group have been completed (S370). If all operations in the group have been completed, metadata corresponding thereto may be updated and flushed (S380). If fewer than all operations in the group have been completed, the next transaction may be processed (S390) (e.g., until all operations in the group have been completed).

[0075] Figure 4 is a block diagram depicting a workflow according to an embodiment of the present disclosure.

[0076] Reference Figure 4, transactions 410 may be received by the system 400 to perform an overwrite merge with transaction grouping, and may be directed to the transaction module 420. In one example, the transaction module 420 may identify a plurality of transactions in a pending queue, the plurality of transactions having one or more key value updates corresponding to a plurality of keys, respectively. The transaction module 420 may identify a common associated key among the plurality of keys that is associated with a common associated key value update among key value updates belonging to different transactions. The transaction module 420 may assign a transaction ID and a transaction group ID to each of the transactions 410. The transaction module 420 may add the most recently received transaction to the most recent group in time. Once the most recent group (or the most recent group, in effect) becomes full, the transaction module 420 may create a new group for subsequent transactions 410. For example, referring to Figure 4 , the transaction module 420 may create a transaction group (eg, Group N) including transactions Trxn1, Trxn2, and Trxn3.

[0077] Return to reference Figure 2 , Group 1 (or Group N) includes transactions Trxn1, Trxn2, and Trxn3, and transactions Trxn1, Trxn2, and Trxn3 include key value updates V1, V2, V3, V4, V5, and V7 corresponding to keys A, B, C, and D, respectively. In this example, key value updates V2 and V3 belonging to transactions Trxn1 and Trxn2, respectively, are associated with key B, and key value updates V4 and V5 belonging to transactions Trxn2 and Trxn3, respectively, are associated with key C. That is, key B is a commonly associated key corresponding to commonly associated key value updates V2 and V3, and key C is a commonly associated key corresponding to commonly associated key value updates V4 and V5.

[0078] After a transaction group (e.g., Group N) is created, a merge module 430 may perform an overwrite merge of key value updates. In one example, the merge module 430 may merge conflicting data writes corresponding to common associated key value updates of a common associated key for transactions grouped in the same transaction group. Then, an inflight request buffer 440 may write the overwrite merged key values ​​to a device (e.g., KVSSD) 450. Once all overwrite merged key values ​​of a transaction group have been successfully written to the device 450, a table 460 containing device metadata may be updated. Finally, the metadata may be written from the table 460 to the device 450.

[0079] According to the above, the embodiments of the present disclosure can group cascaded transactions by using inline overwrite merging, so that the metadata write operation can be delayed until all operations in the associated transaction group are written to the corresponding device, and the associated clearing of metadata can be delayed. Therefore, the embodiments of the present disclosure can achieve crash recovery.

[0080] In addition, embodiments of the present disclosure may be intended to form a transaction group with the maximum possible number of cascaded transactions while avoiding undue delays associated with metadata write operations, and at the same time also dividing non-cascaded transactions into different corresponding transaction groups. Therefore, the implementation of transaction groups according to the disclosed embodiments enables overwrite merging and transactions to work seamlessly while achieving crash consistency and reducing suspended transactions, thereby improving data storage technology.

Claims

1. A method for data storage, the method comprising: identifying a plurality of transactions in a pending queue, the plurality of transactions having one or more key value updates corresponding to a plurality of keys, respectively; identifying a common associated key among the plurality of keys, the common associated key being associated with a common associated key value update among key value updates belonging to different transactions among the plurality of transactions; Based on the corresponding transaction IDs assigned to the transactions, assigning transaction group IDs to the plurality of transactions respectively; grouping the plurality of transactions into corresponding transaction groups among a plurality of transaction groups based on the assigned transaction group ID; For grouped transactions in the same transaction group in the multiple transaction groups, merging conflicting data writes corresponding to updates of common association key values ​​of common association keys; writing all merged key value updates written to conflicting data of transaction groups in the plurality of transaction groups to one or more storage devices; as well as Only when all the merged key value updates are confirmed to have been written to the one or more storage devices, metadata corresponding to the merged key value updates of the transaction group in the plurality of transaction groups is updated.

2. The method according to claim 1, wherein: When a crash occurs at a crashed transaction among transactions in a crashed transaction group among the plurality of transaction groups, the method further includes retrying only transactions among transactions in the crashed transaction group that are linked to the crashed transaction due to merging of conflicting data writes.

3. The method of claim 1, further comprising: Removing transactions from the plurality of transactions that have only one or more earlier key-value updates in time, the one or more earlier key-value updates being unrelated to one or more later key-value updates corresponding to the same respective key, the later key-value updates being in one or more other transactions from the plurality of transactions that occur later in time than the removed transaction from the plurality of transactions.

4. The method of claim 1, further comprising: determining a pair of consecutive transactions lacking any commonly associated key value updates corresponding to a common key among the plurality of keys; as well as Transaction group IDs are respectively allocated to the consecutive transactions, so that the consecutive transactions are located in different corresponding transaction groups.

5. The method of claim 1, further comprising: determining which pair of consecutive transactions has a minimum number of commonly associated key value updates corresponding to a common key in the plurality of keys; as well as Transaction group IDs are respectively allocated to the consecutive transactions, so that the consecutive transactions are located in different corresponding transaction groups.

6. The method according to claim 1, wherein: Assigning a transaction group ID is also based on an analysis of the total number of commonly associated key value updates across different transactions.

7. The method according to claim 1, further comprising: Transaction IDs are assigned to the plurality of transactions.

8. A system for performing a merge with transaction grouping, the system comprising: Transaction module, used to: identifying a plurality of transactions in a pending queue, the plurality of transactions having one or more key-value updates corresponding to a plurality of keys, respectively, identifying a common association key of the plurality of keys that is associated with a common association key value update of key value updates belonging to different transactions of the plurality of transactions, assigning transaction group IDs to the plurality of transactions, respectively, based on the corresponding transaction IDs assigned to the transactions, and grouping the plurality of transactions into corresponding transaction groups among a plurality of transaction groups based on the assigned transaction group ID; a merging module, configured to merge conflicting data writes corresponding to updates of common association key values ​​of common association keys for grouped transactions in the same transaction group in the multiple transaction groups; as well as The in-flight request buffer is configured to: write all merged key value updates of conflicting data written by a transaction group among the multiple transaction groups to one or more storage devices, and update metadata corresponding to the merged key value updates of the transaction group among the multiple transaction groups only when it is confirmed that all merged key value updates have been written to the one or more storage devices.

9. The system according to claim 8, wherein: When a crash occurs at a crashed transaction in transactions in a crashed transaction group in the plurality of transaction groups, the transaction module is further configured to retry only transactions in the crashed transaction group that are linked to the crashed transaction due to merging conflicting data writes.

10. The system according to claim 8 or 9, wherein: The merge module is also configured to remove transactions from the multiple transactions that only have one or more earlier key value updates, wherein the one or more earlier key value updates are unrelated to one or more later key value updates corresponding to the same respective key, wherein the later key value updates are in one or more other transactions from the multiple transactions that occur later in time than the removed transaction from the multiple transactions.

11. The system according to claim 8 or 9, wherein: The transaction module is also configured to: determining a pair of consecutive transactions lacking any commonly associated key value updates corresponding to a common key of the plurality of keys; and Transaction group IDs are respectively allocated to the consecutive transactions, so that the consecutive transactions are located in different corresponding transaction groups.

12. The system according to claim 8 or 9, wherein: The transaction module is also configured to: determining which pair of consecutive transactions has a minimum number of commonly associated key value updates corresponding to a common key in the plurality of keys; and Transaction group IDs are respectively allocated to the consecutive transactions, so that the consecutive transactions are located in different corresponding transaction groups.

13. The system according to claim 8 or 9, wherein: The transaction module is further configured to assign a transaction group ID based on an analysis of a total number of commonly associated key value updates across different transactions.

14. The system according to claim 8 or 9, wherein: The transaction module is further configured to assign transaction IDs to the plurality of transactions.

15. A computer-readable storage medium storing a program, which, when executed by a processor, causes the processor to execute the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Data integrity and loss resistance in high performance and high capacity storage deduplication

    CN105843551A

  • Data batch processing method and device

    CN106844507A