A Method for Reducing Write Amplification and Write Stall of LSM Tree Based on NVM
By fine-graining the compression of L0 to L1 layers in the LSM tree and using NVM to expand capacity, the write amplification and write stall problems of LSM tree are solved, and more efficient data storage performance is achieved.
Patent Information
- Application Number
- CN202411343492.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-25
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2044-09-25
AI Technical Summary
The existing LSM trees have write amplification and write stall problems in write-intensive scenarios, especially in compression operations from L0 to L1, which leads to performance degradation.
The compression of L0 to L1 layers in the LSM tree is fine-grained, and the capacity of each layer is expanded through NVM (nonvolatile memory), changing the storage structure of L1 layer, reducing write amplification and write stalling.
It reduces write amplification and write stall during LSM tree compression, improves read and write efficiency, and improves overall performance with the fast read performance of NVM.
Smart Images

Figure CN119376618B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of information processing technology, and particularly relates to a method for reducing write amplification and write stall of an LSM tree based on NVM. Background Art
[0002] Since the emergence of key-value storage technology, key-value storage systems have played a crucial role in large-scale, high-performance, data-intensive application scenarios, and the LSM tree has played a key role as its core data structure. Thanks to write optimization, the LSM tree has excellent write performance and is very suitable for dealing with write-intensive scenarios, and the structure of hierarchical sorting of data also retains good read performance for it.
[0003] However, the current LSM architecture uses SSTable as its basic compression unit, and a large amount of data is involved at one time in the compression operation from L0 layer to L1 layer, so there are problems of high write amplification and write stall. The present invention discloses a method for reducing write amplification and write stall of an LSM tree based on NVM. Summary of the Invention
[0004] To solve the problems existing in the prior art, the present invention provides a method for reducing write amplification and write stall of an LSM tree based on NVM. The method of the present invention makes the compression between L0 layer and L1 layer of the LSM tree fine-grained, expands the capacity of each layer of the LSM tree in equal proportion, and changes the storage structure of the L1 layer. After that, it can reduce the write stall and write amplification in the compression process of the LSM tree, improve the read and write efficiency, and solve the problems mentioned in the above background art.
[0005] To achieve the above object, the present invention provides the following technical solution: A method for reducing write amplification and write stall of an LSM tree based on NVM, including the following steps:
[0006] S1. Initialize the storage system;
[0007] S2. Accept user input;
[0008] S3. If the accepted operation is a read operation, execute step S4; otherwise, execute step S5;
[0009] S4. Query the target data in sequence according to the order of MemTable, Immutable MemTable, and L0 to Ln layers until the search is successful;
[0010] S5. If the accepted operation is a write operation, execute step S6; otherwise, execute step S7;
[0011] S6. The system writes the data into the MemTable in the memory;
[0012] S7. If the MemTable is full, execute step S8; otherwise, execute step S9;
[0013] S8. Convert the MemTable into an Immutable MemTable, and generate metadata from the Immutable MemTable and insert it into the metadata group of level L0 when the system is idle;
[0014] S9. When the amount of data contained in level L0 reaches the threshold, execute step S10; otherwise, execute step S11;
[0015] S10. Select multiple key ranges from the metadata group of level L0, and logically compress all key-value pairs contained therein into the complementary table of the corresponding Block table;
[0016] S11. When the amount of data within the key range in the metadata group of level L0 reaches the threshold, execute step S12; otherwise, execute step S13;
[0017] S12. Compress all data within this key range to level L1;
[0018] S13. If the Block table is too large, execute step S14; otherwise, execute step S15;
[0019] S14. Split the over-large Block table into multiple Block tables and change the key ranges in the metadata group of level L0;
[0020] S15. If there exists a layer Li (n > i > 0) whose internal contained data volume reaches the threshold, execute step S16; otherwise, execute step S17;
[0021] S16. Compress data from layer Li (n > i > 0) to layer Li+1, and execute step S15;
[0022] S17. When the life cycle of the storage system ends, the key-value storage program is completed.
[0023] Preferably, in step S1, the storage system includes a memory part, an NVM part, and an SSD part; the memory part includes a MemTable and an Immutable MemTable; the NVM part contains a log, level L0, and layer L1+, where layer L1+ represents the part of layer L1 in the NVM; the SSD part includes layers L1 to Ln, where layer Ln is the deepest layer that the LSM tree can reach.
[0024] Preferably, the L0 layer is organized into multiple Log tables, each Log table is directly converted from Log, and the key-value pairs in each Log table are logically sorted by the metadata in its metadata group. The metadata group is divided into multiple key ranges, where each key range corresponds to and is the same as the key range of each Block table in the L1 layer, and each key range contains multiple metadata;
[0025] The L1 layer located in the SSD consists of multiple Block groups, each Block group contains multiple blocks for storing ordered key-value pairs; the L1+ layer located in the NVM consists of multiple supplementary tables, and a supplementary table is also used to store ordered key-value pairs. Each supplementary table corresponds to a Block group one by one, and a pair of Block group and supplementary table forms a Block table.
[0026] Preferably, in step S4, during reading, first query the MemTable and ImmutableMemTable in the memory. If not hit, continue to query the L0 to Ln layers in sequence downwards;
[0027] When querying the L0 layer, first confirm a key range in the metadata group through binary search, and then filter and query the target data through the Bloom filter of each piece of metadata therein;
[0028] When querying the L1 layer, first determine the Block table where the target key-value pair is located through binary search, and then query the target data from the Block group and the supplementary table through the metadata.
[0029] Preferably, in step S8, when each Immutable MemTable generates metadata, it needs to generate multiple pieces of metadata separately according to the key ranges in the metadata group of the L0 layer, and then insert each piece of metadata into the metadata group in the L0 layer according to the key range it belongs to. At the same time, convert the corresponding Log logically into a Log table and insert it into the L0 layer.
[0030] Preferably, the Block table includes three parts: key-value data in the SSD, key-value data in the NVM, and metadata; among them, the metadata and the key-value data in the NVM are jointly stored in the supplementary table, while the key-value data in the SSD is organized into blocks and stored in the Block group.
[0031] Preferably, in step S12, the key space of the L1 layer is divided into multiple continuous key ranges by the minimum key of all Block tables. Among them, the first two key ranges are the key ranges of the first Block table, and each subsequent key range corresponds to each subsequent Block table in sequence; if the L1 layer is empty, it is regarded as an empty Block table that only contains an infinite key range.
[0032] The beneficial effects of the present invention are as follows:
[0033] 1) Less write amplification: A large amount of write amplification is generated during the compression of the LSM tree. The present invention realizes an LSM tree with a lower number of levels and a delayed compression L1 layer, effectively reducing the write amplification brought by the compression operation;
[0034] 2) Less write stall: Usually, a large amount of data is involved at one time when compressing from the L0 layer to the L1 layer. The present invention realizes the decomposition of the compression from the L0 layer to the L1 layer, reducing the tail latency during the operation of the LSM tree;
[0035] 3) Faster read speed: NVM has a faster read speed than SSD. The present invention realizes an LSM tree with a faster read speed by using NVM. Description of the Drawings
[0036] Figure 1 It is the architecture diagram of the method for reducing write amplification and write stall of the LSM tree based on NVM in the embodiment of the present invention;
[0037] Figure 2 It is the data storage structure diagram of the Block table in the embodiment of the present invention;
[0038] Figure 3 It is the flowchart of the method for reducing write amplification and write stall of the LSM tree based on NVM in the embodiment of the present invention. Detailed Embodiments
[0039] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0040] Please refer to Figures 1 - 3 , the present invention provides a technical solution: a method for reducing write amplification and write stall of the LSM tree based on NVM. After fine-graining the compression between the L0 layer and the L1 layer in the LSM tree, expanding the capacity of each layer of the LSM tree in equal proportion, and changing the storage structure of the L1 layer, the write stall and write amplification during the compression of the LSM tree can be reduced. The method includes several parts such as the overall structure, reading, writing, and compression.
[0041] Overall structure: The storage system of the present invention includes a memory part, an NVM part, and an SSD part. As Figure 1As shown in the figure, the memory part includes a MemTable and an Immutable MemTable; the NVM part includes a log, an L0 layer, and an L1+ layer, where the L1+ layer represents the part of the L1 layer in the NVM; the SSD part includes L1 to Ln layers, where the Ln layer is the deepest layer that the LSM tree can reach in the present invention. The data capacity of each layer after modification in the present invention is much higher than that of each layer in the traditional LSM tree, and the number of layers of the LSM tree is reduced.
[0042] The design of the memory part is the same as that of the traditional LSM tree.
[0043] The L0 layer is organized into multiple Log tables, each of which is directly converted from a Log without modification. The key-value pairs in each Log table are logically sorted by the metadata in its metadata group. The metadata group is divided into multiple key ranges, where each key range corresponds to and is the same as the key range of each Block table in the L1 layer, and each key range contains multiple metadata.
[0044] The L1 layer located in the SSD consists of multiple Block groups, and each Block group contains multiple blocks for storing ordered key-value pairs. The L1+ layer located in the NVM consists of multiple supplementary tables, and one supplementary table is also used to store ordered key-value pairs. Each supplementary table corresponds to a Block group one by one, and a pair of Block group and supplementary table forms a Block table. As Figure 2 shown in the figure, the Block table includes three parts: key-value data in the SSD, key-value data in the NVM, and metadata. The metadata and the key-value data in the NVM are stored together in the supplementary table, while the key-value data in the SSD is organized into blocks and stored in the Block group. All Block tables are ordered and non-overlapping among tables, and each block within a Block table is also ordered and non-overlapping. Moreover, the key-value pairs in the Block table are logically sorted by the metadata. The key space of the L1 layer is divided into multiple consecutive key ranges by the minimum key of all Block tables. The first two key ranges are the key ranges of the first Block table, and each subsequent key range corresponds to each subsequent Block table in sequence. If the L1 layer is empty, it is regarded as an empty Block table that only contains an infinite key range.
[0045] Finally, the Li (n≥i>1) layer has the same layout as the traditional LSM tree and is stored in the SSD in the format of an SSTable.
[0046] Reading: When reading, first query the MemTable and Immutable MemTable in the memory. If not hit, continue to query the L0 to Ln layers sequentially downward.
[0047] When querying the L0 layer, first confirm a key range in the metadata group through binary search, and then filter and query the target data through the Bloom filter of each piece of metadata. When querying the L1 layer, first determine the Block table where the target key-value pair is located through binary search, and then query the target data from the Block group and the supplementary table through the metadata. The query method for the Li layer (n≥i>1) is the same as that of the traditional LSM tree.
[0048] Writing and Compression: When writing, the data is first written into the MemTable and the Log. When the MemTable is full, it is converted into an Immutable MemTable and used to generate metadata. When each Immutable MemTable generates metadata, multiple pieces of metadata need to be generated separately according to the key range in the L0 layer metadata group, and then each piece of metadata is inserted into the metadata group in the L0 layer according to the belonging key range. At the same time, the corresponding Log logic is converted into a Log table and inserted into the L0 layer.
[0049] When the data volume in the L0 layer reaches the threshold, it triggers the data compression from the L0 layer to the L1 layer. At this time, multiple key ranges are selected from the L0 layer metadata group, and all the key-value pairs contained therein are logically compressed into the supplementary table of the corresponding Block table, and the compression and sorting are performed by updating the metadata in the Block table. If the Block table is too large, the too-large Block table is split into multiple Block tables, and the key range in the L0 layer metadata group is changed. When the data volume within the Key range in the L0 layer metadata group reaches the threshold, the compression operation from the L0 layer to the L1 layer is also triggered, and all the data within this key range is compressed to the L1 layer.
[0050] When the data volume in the L1 layer reaches the threshold, it triggers the data compression from the L1 layer to the L2 layer. During compression, first select multiple Block tables from the L1 layer, then select the SSTables overlapping with these Block tables from the L2 layer, read and sort the selected tables together, and finally organize them into SSTables and write them back to the L2 layer, and change the key range in the L0 layer metadata group.
[0051] The compression of the Li layer (n>i>1) is the same as that of the traditional LSM tree.
[0052] A method for reducing write amplification and write stall of an LSM tree based on NVM in the present invention, as Figure 3 shown, includes the following implementation steps:
[0053] (1) Initialize the storage system;
[0054] (2) Accept user input;
[0055] (3) If a read operation is accepted, execute step (4); otherwise, execute step (5);
[0056] (4) Query the target data in the order of MemTable, Immutable MemTable, and L0 to Ln layers in sequence until the search is successful;
[0057] (5) If a write operation is accepted, execute step (6); otherwise, execute step (7);
[0058] (6) The system writes the data into the MemTable in the memory;
[0059] (7) If the MemTable is full, execute step (8); otherwise, execute step (9);
[0060] (8) Convert the MemTable into an Immutable MemTable, and generate metadata through the Immutable MemTable and insert it into the L0 layer metadata group when the system is idle;
[0061] (9) When the amount of data contained in the L0 layer reaches the threshold, execute step (10); otherwise, execute step (11);
[0062] (10) Select multiple key ranges from the L0 layer metadata group, and logically compress all the key-value pairs contained therein into the supplementary table of the corresponding Block table;
[0063] (11) When the amount of data within the key range in the L0 layer metadata group reaches the threshold, execute step (12); otherwise, execute step (13);
[0064] (12) Compress all the data within this key range to the L1 layer;
[0065] (13) If the Block table is too large, execute step (14); otherwise, execute step (15);
[0066] (14) Split the overly large Block table into multiple Block tables, and change the key range in the L0 layer metadata group;
[0067] (15) If there exists an Li (n > i > 0) layer where the amount of data contained therein reaches the threshold, execute step (16); otherwise, execute step (17);
[0068] (16) Compress the data from the Li (n > i > 0) layer to the Li+1 layer, and execute step (15);
[0069] (17) If the storage system life cycle ends, the key-value storage program is completed; otherwise, return to step (2).
[0070] So far, all the implementation steps of the method for reducing write amplification and write stall of the LSM tree based on NVM according to the present invention are completed.
[0071] In the embodiments provided by the embodiments of the present invention, it should be noted that in this article, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, article or device including the said element.
[0072] The terms used in the embodiments of the present invention are only for the purpose of describing specific embodiments, and are not intended to limit the present invention. The singular forms "a", "the" and "said" used in the embodiments of the present invention and the appended claims are also intended to include the plural forms, unless the context clearly indicates otherwise.
[0073] It should be understood that the term "and / or" used herein is only a kind of association relationship describing associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this article generally represents an "or" relationship between the front and rear associated objects.
[0074] Depending on the context, the word "if" as used herein can be interpreted as "when", "while", "in response to a determination" or "in response to a detection". Similarly, depending on the context, the phrase "if a determination" or "if a detection (stated condition or event)" can be interpreted as "when a determination is made", "in response to a determination", "when a detection (stated condition or event) is made" or "in response to a detection (stated condition or event)".
[0075] The "first / second" mentioned in the embodiments is only to distinguish similar objects and does not represent a specific order for the objects. It can be understood that the "first / second" can be interchanged in a specific order or sequence when allowed. It should be understood that the objects distinguished by the "first / second" can be interchanged appropriately so that the embodiments described here can be implemented in an order other than those illustrated or described here.
[0076] Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for reducing write amplification and write stall of an LSM tree based on NVM, characterized in that It includes the following steps: S1. Initialize the storage system; the storage system includes a memory part, an NVM part, and an SSD part; the memory part includes a MemTable and an Immutable MemTable; the NVM part contains a log, an L0 layer, and an L1+ layer, where the L1+ layer represents the part of the L1 layer in the NVM; the SSD part includes L1 to Ln layers, where the Ln layer is the deepest layer that the LSM tree can reach; The L0 layer is organized into multiple Log tables, each of which is directly converted from a Log. The key-value pairs in each Log table are logically sorted by their metadata in the metadata group. The metadata group is divided into multiple key ranges, where each key range corresponds to and is the same as the key range of each Block table in the L1 layer, and each key range contains multiple metadata; The L1 layer located in the SSD consists of multiple Block groups, and each Block group contains multiple blocks for storing ordered key-value pairs; The L1+ layer located in the NVM consists of multiple supplementary tables. A supplementary table is also used to store ordered key-value pairs. Each supplementary table corresponds to a Block group one by one. A pair of Block groups and a supplementary table form a Block table; S2. Accept user input; S3. If a read operation is accepted, execute step S4; otherwise, execute step S5; S4. Query the target data in the order of the MemTable, the Immutable MemTable, and the L0 to Ln layers in sequence until the search is successful; S5. If a write operation is accepted, execute step S6; otherwise, execute step S7; S6. The system writes the data into the MemTable in the memory; S7. If the MemTable is full, execute step S8; Otherwise, execute step S9; S8. Convert the MemTable into an Immutable MemTable, and generate metadata through the Immutable MemTable and insert it into the metadata group of the L0 layer when the system is idle; S9. When the data volume contained in the L0 layer reaches the threshold, execute step S10; Otherwise, execute step S11; S10. Select multiple key ranges from the metadata group of the L0 layer, and logically compress all the key-value pairs contained therein into the supplementary tables of the corresponding Block tables; S11. When the data volume within the key range in the metadata group of the L0 layer reaches the threshold, execute step S12; Otherwise, execute step S13; S12. Compress all the data within this key range to the L1 layer; the key space of the L1 layer is divided into multiple consecutive key ranges by the minimum key of all Block tables. The first two key ranges are the key ranges of the first Block table, and each subsequent key range corresponds to each subsequent Block table in sequence; If the L1 layer is empty, it is regarded as an empty Block table that only contains an infinite key range; S13. If the Block table is too large, execute step S14; Otherwise, execute step S15; S14. Split an overly large Block table into multiple Block tables, and change the key range in the metadata group of level L0; S15. If there exists a level Li (n > i > 0) whose internal data volume reaches the threshold, execute step S16; otherwise, execute step S17; S16. Compress data from level Li (n > i > 0) to level Li+1, and execute step S15; S17. When the storage system life cycle ends, the key-value storage program is completed.
2. The method for reducing write amplification and write stall of the LSM tree based on NVM according to claim 1, wherein: In step S4, during reading, first query the MemTable and Immutable MemTable in the memory. If not hit, continue to query levels L0 to Ln sequentially downward; When querying level L0, first confirm a key range in the metadata group through binary search, and then filter and query the target data through the Bloom filter of each piece of metadata therein; When querying level L1, first determine the Block table where the target key-value pair is located through binary search, and then query the target data from the Block group and the supplementary table through the metadata; 3. The method for reducing write amplification and write stall of the LSM tree based on NVM according to claim 1, wherein: In step S8, when each Immutable MemTable generates metadata, it needs to generate multiple pieces of metadata separately according to the key range in the metadata group of level L0, and then insert each piece of metadata into the metadata group in level L0 according to the belonging key range. At the same time, convert the corresponding Log logic into a Log table and insert it into level L0.
4. The method for reducing write amplification and write stall of the LSM tree based on NVM according to claim 1, wherein: The Block table includes three parts: key-value data in the SSD, key-value data in the NVM, and metadata; among them, the metadata and the key-value data in the NVM are jointly stored in the supplementary table, while the key-value data in the SSD is organized into blocks and stored in the Block group.
Citation Information
Patent Citations
Key value storage system based on NVM and SSD hybrid storage structure
CN110347336A
Cache optimization method for reading performance of KV storage system based on LSM-tree
CN114398007A
Key value storage write optimization method based on nonvolatile memory
CN117908755A