An Early Compaction Method and System Applied to the LSM Tree Structure
By dividing the task queue according to the overlap of key ranges of SSTable in the LSM tree structure and using file granular pipeline optimization, the read and write amplification problem caused by Compaction operations in the LSM tree structure is solved, and the system performance is improved.
Patent Information
- Application Number
- CN202510346074.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-03-24
AI Technical Summary
Compaction operations of the existing LSM tree structures cause read and write amplification problems when facing a large amount of data, especially when there is scope overlap in continuous merge operations of adjacent layers, resulting in system performance degradation.
The advance Compaction method is adopted to divide the task queue into two operation queues according to the overlap of key ranges of SSTable, only modify the metadata or perform large-scale Compaction operations, and speed up the write back and generation process through file granular pipeline optimization scheme.
It effectively alleviates the problem of repeated read and write amplification due to range overlap in continuous Compaction operations, improves the efficiency of Compaction operations, reduces the total operating time, and improves system performance.
Smart Images

Figure CN119861880B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer storage, and particularly to an early Compaction method and system applied to the LSM tree structure. Background Art
[0002] With the advent of the digital age, people have increasingly paid attention to data storage and management. The traditional storage architecture is not obvious in the face of a large amount of high-frequency data writing. In this case, the key-value storage system based on the LSM tree has emerged. As an innovative data structure, the emergence of the LSM tree structure has proposed an effective solution to the problems existing in traditional data storage.
[0003] The LSM tree uses an append-only storage method to implement sequential write operations. Due to this feature, the LSM tree provides excellent write performance. The main components of the LSM tree include the log file on the disk, the sorted string table SSTable on the disk, and the variable MemTable in the memory. When new write and update requests arrive, they are first written into the log file on the disk and then into the MemTable in the memory. When the current MemTable capacity reaches its threshold, a new MemTable and log file will be created to store subsequent write and update requests. In the background, the previous variable MemTable will be converted into an immutable Immutable MemTable, and then the compression thread will flush it to the disk, and a new SSTable will be generated in the L0 layer of the LSM tree. At this time, the previous log file can be discarded. When the data flushed from the memory to the L0 layer reaches the threshold of the data volume that the L0 layer can store, the SSTable in the L0 layer will be flushed and merged into the L1 layer through the Compaction operation. The specific process of the Compaction operation is as follows: When the data storage volume of the L i layer reaches its threshold, first select an SSTable from the L i layer, and then select all SSTables in the L i+1 layer whose key ranges intersect with the key range of the SSTable in the L i layer. Then read them into the memory for merge sorting to generate a new file, and finally write it back to the disk.
[0004] Although the global order among SSTables can be maintained during the Compaction operation of the LSM tree, frequent rewrite operations during this period can cause the problem of read-write amplification. When the data volume is extremely large, the Compaction operation may also lead to the occurrence of write stalls, resulting in a decrease in system performance. Currently, many studies have also proposed some solutions, such as using key-value separation to reduce the amount of data read during the Compaction operation. However, this solution does not consider the association between consecutive merge operations in adjacent layers. During the actual operation process, for consecutive merge operations in adjacent layers, if the key range selected by layer L i each time overlaps with the key range of layer L i+1 then the data in the overlapping part will be read into memory multiple times for merge sorting and then written back to disk, resulting in unnecessary I / O amplification.
[0005] In summary, a strategy is needed that can not only keep the SSTables in an ordered state during the Compaction operation but also alleviate the additional read-write amplification problem caused by range overlap during consecutive Compaction operations. Summary of the Invention
[0006] The purpose of the present invention is to solve the problems in the prior art.
[0007] The technical solution adopted by the present invention to solve its technical problems is: to provide an early Compaction method applied to the LSM tree structure, including the following steps:
[0008] When the data volume stored in the i-th layer L i of the LSM tree reaches its threshold, trigger the Compaction operation;
[0009] Select the SSTable read tasks of multiple consecutive Compaction operations to be executed subsequently in layer L i+1 and layer L i and put them into the task queue;
[0010] After reading multiple SSTables into the task queue, judge whether to perform a large Compaction operation or only modify the metadata according to different key range overlap situations to implement the request for the movement of SSTables between layer L i and layer L i+1 .
[0011] Preferably, the judging whether to perform a large Compaction operation or only modify the metadata according to different key range overlap situations to implement the movement of SSTables between layer L i and layer L i+1A request for movement between layers, including the following steps:
[0012] Divide the task queue into two operation queues according to whether there is a key range overlap in the SSTable. If there is no range overlap between the SSTables in layer L i and layer L i+1 , divide it into the first operation queue; if there is a range overlap between the SSTables in layer L i and layer L i+1 , divide it into the second operation queue;
[0013] For the first operation queue, only modify the metadata of layer L i and layer L i+1 , and then insert the SSTable of layer L i into the corresponding position of layer L i+1 ; for the second operation queue, merge multiple subsequent consecutive Compaction operation tasks with the current Compaction operation task into a large Compaction operation task and execute them together to realize the migration of SSTable from layer L i and layer L i+1 .
[0014] Preferably, there are three cases for the second operation queue, and different treatments are carried out according to different cases, including:
[0015] For the intersection and non-intersection relationships, find the position where the newly generated SSTable after the Compaction operation task is stored in layer L i+1 so that it is globally ordered in layer L i+1 ; among them, the SSTable that has a range overlap with multiple Compaction operations at the same time only needs to be added to the queue once;
[0016] For the disjoint relationship, set an SSTable interval stop flag. When generating a new SSTable, the SSTable will no longer continue to be written before the minimum key of the interval SSTable, and then create a new SSTable to accept the subsequent data writing.
[0017] Preferably, for the second operation queue, in order to speed up the write-back speed of the SSTable, the write-back of the previous SSTable file and the generation of the next SSTable file are carried out simultaneously.
[0018] Preferably, the simultaneous write-back of the previous SSTable file and the generation of the next SSTable file are implemented by using a file-level pipeline scheme, which specifically includes the following steps:
[0019] Creation step: According to the task information returned in the task queue, create an iterator for the corresponding SSTable file in this task for subsequent traversal of relevant data, and create a file data structure of the SSTable in memory for receiving the subsequent generated ordered data blocks;
[0020] Writing step: Traverse the data according to the iterator, delete the expired key-value pair information, retain the valid data, and organize the valid data into multiple data blocks and write them into the created SSTable file data structure;
[0021] When the capacity of the data blocks stored in this SSTable reaches its threshold, add the corresponding metadata to this SSTable and write it back to the L i+1 layer; During the process of writing the previous SSTable back to disk, the iterator also continues to traverse the valid data and create a new SSTable file data structure in memory to receive the new valid data blocks; When all the valid data is written back to the L i+1 layer in an orderly manner, the operation of early Compaction is completed.
[0022] The present invention also provides an early Compaction system applied to the LSM tree structure for implementing the method described in any one of the above, including:
[0023] Triggering module: When the data volume stored in the i-th layer L of the LSM tree i reaches its threshold, trigger the Compaction operation;
[0024] Partitioning module: Select the L of the i+1-th layer i+1 and the L i layer, and read the SSTables of multiple consecutive Compaction operations to be executed later into the task queue;
[0025] Compaction module: After reading multiple SSTables into the task queue, judge whether to perform a large Compaction operation or only modify the metadata according to the overlap situation of different key ranges to implement the request for the movement of the SSTable between the L i layer and the L i+1 layer.
[0026] The present invention has the following beneficial effects:
[0027] (1) The present invention effectively alleviates the I / O amplification problem of repeated reading and writing caused by the overlap of ranges in consecutive Compaction operations. Through the present invention, the overlapping part of the ranges only needs to be read into memory once and written back to disk once to complete the Compaction operation.
[0028] (2) The present invention accelerates the execution efficiency of the Compaction operation by using the optimized scheme of the file - granularity pipeline. On the one hand, since the early Compaction strategy may cause the write - back and generation of SSTables in memory to take a relatively long time, using this optimized scheme enables the write - back and generation operations to be executed in parallel in a pipeline manner, and ultimately the total time actually spent is only determined by the time spent on the SSTable generation operation. On the other hand, this optimized scheme can be used regardless of the number of SSTables or the level to which the final SSTable is to be written back, that is, this optimized scheme has general applicability.
[0029] The following further elaborates on the present invention in detail with reference to the accompanying drawings and embodiments, but the present invention is not limited to the embodiments. Description of the Drawings
[0030] Figure 1 It is the method - step diagram of the embodiment of the present invention;
[0031] Figure 2 It is the detailed flowchart of the embodiment of the present invention;
[0032] Figure 3 It is the schematic diagram of the second operation queue situation of the embodiment of the present invention;
[0033] Figure 4 It is the schematic diagram of the file - granularity pipeline scheme of the embodiment of the present invention;
[0034] Figure 5 It is the system structure diagram of the embodiment of the present invention. Specific Embodiments
[0035] The early Compaction scheme proposed by the present invention applied to the LSM - tree structure: In the LSM - tree, when the L i layer reaches the threshold of the SSTable storage capacity, the Compaction operation is triggered. After being triggered, the Compaction operation is not directly performed, but the SSTables to be subjected to the Compaction operation later are taken. These Compaction operations are added to the task queue of the early Compaction operation. When the length of the task queue reaches its threshold or there is no overlapping situation in the SSTable range between the L i layer and the L i+1 layer, this operation is no longer performed. Immediately afterwards, the task queue needs to be processed. Since there may be two situations in the key - overlapping range between SSTables, the task queue needs to be further divided into two operation queues for processing. The specific operations are as follows: (i) The first operation queue: When the L i layer and the Li+1 When there is no range overlap between SSTables in the layer, directly modify L i+1 the layer-level information metadata, and then insert the SSTable into L i+1 at the corresponding position of the layer to make the layer globally ordered; (ii) The second operation queue: When there is a range overlap between the SSTables in layer L i and layer L i+1 combine multiple Compaction operations in the queue into a large Compaction operation and perform them simultaneously. This large Compaction operation is the process of merging and sorting all the SSTables in the queue, and then writing them back to the disk. It should be noted that when writing the merged and sorted SSTables back to the disk, file-level pipelining is used for write-back optimization. Based on the above LSM tree pre-Compaction scheme, the components and designs of the present invention are: a Compaction task queue. This queue is implemented using a linked list structure, and each node in the linked list stores a Compaction task. In other words, each node stores the SSTable files corresponding to layer L i and layer L i+1 After the task queue is obtained, it needs to be divided into two specific operation queues according to different situations of the SSTables. Multiple consecutive Compaction tasks with range overlap will be combined into a large Compaction task, and then the corresponding Compaction task information and the request to write the SSTable back to the disk will be returned.
[0036] Specifically, refer to Figure 1 and Figure 2 as shown, which is the method step diagram and detailed flowchart of the embodiment of the present invention, including the following steps:
[0037] S101, when the data volume stored in the i-th layer L of the LSM tree i reaches its threshold, trigger a Compaction operation;
[0038] S102, select the i + 1-th layer L i+1 and read the SSTables of multiple consecutive Compaction operations to be executed subsequently in layer L i into the task queue;
[0039] S103, after reading multiple SSTables into the task queue, determine whether to perform a large Compaction operation or only modify the metadata according to different key range overlap situations to implement the SSTable in layer L i and layer L i+1Request to move between layers.
[0040] As Figure 3 shown, there may be three cases for the second operation queue: intersection relationship, non - intersection relationship, and separation relationship. For the intersection and non - intersection relationships, only the position where the newly generated SSTable after the Compaction operation task is executed is stored in the L i+1 layer needs to be found so that the L i+1 layer is globally ordered. It should be noted that for an SSTable that overlaps in range with multiple Compaction operations, it only needs to be added to the queue once. For the separation relationship, since there are SSTable files that do not participate in the early Compaction operation between the SSTable files selected for the early Compaction operation in the L i+1 layer, when the new SSTable is written back to the L i+1 layer after the early Compaction operation is completed, the new SSTable may overlap in range with the SSTable files that do not participate in the early Compaction operation, which may cause the L i+1 layer to no longer maintain the state of global order or even data loss. To solve this problem, an SSTable interval stop flag is set. When generating a new SSTable, this SSTable is no longer written before the minimum key of the interval SSTable, and then a new SSTable is created to accept the subsequent data writes. According to the range of the Compaction operation, the subsequent data must be greater than the maximum key of this interval SSTable. Thus, we ensure that the L i+1 layer remains globally ordered in the case of separation.
[0041] Specifically, taking the intersection case as an example to illustrate the process of the embodiment of the present invention, as Figure 2 shown. When the SSTable1 in the L i layer needs to perform a Compaction operation with the SSTable3 and SSTable4 in the L i+1 layer, this Compaction task is put into the queue, as shown in step 1 of Figure 2 ; then continue to search and find that the next Compaction operation needed is for SSTable2, SSTable4, and SSTable5. So this Compaction task is also put into the next node in the queue, as shown in Figure 2As shown in Step 2. It can be found that at this time, for SSTable4 in this intersection case, only one read and one write-back are required to complete the Compaction operation. After all overlapping tasks are put into the task queue or the length of the task queue reaches its threshold, all Compaction tasks are merged into a large Compaction task for operation, that is, all valid data is merged and sorted to form data blocks, as Figure 2 shown in Step 3. Then the generated data blocks are written into the SSTable file structure created before, as Figure 2 shown in Step 4. Finally, after adding the corresponding SSTable file metadata, it is written back to the disk, as Figure 2 shown in Step 5-1. During the process of writing SSTable6 back to the disk, the generation of the next data block is also in progress, as Figure 2 shown in Step 5-2. Figure 2 There are two Step 5s to illustrate that these two steps are executed in parallel.
[0042] See Figure 4 shown. It is a schematic diagram of the file granularity pipeline optimization scheme of the present invention, including the following steps: During the execution of the Compaction operation, since all the written-back SSTables are independent of each other, and since the time spent on writing back the SSTable and generating the SSTable belongs to the same order of magnitude, the write-back of the previous SSTable file and the generation of the next SSTable file can be carried out at the same time. Assume that the time for the SSTable generation part is t1, the time for the SSTable write-back part is t2, and there are a total of n SSTables to be operated in the task queue. From Figure 4 it can be easily obtained that the total time required to complete the early Compaction operation under this optimization scheme is n*t1 + t2, and if this optimization scheme is not adopted, the total time required to complete is n*t1 + n*t2. Through the calculation of the above total time, it can be seen that adopting this optimization scheme can effectively reduce the operation time and improve the system operation efficiency.
[0043] To verify the effectiveness of the present invention, verification experiments were conducted. In the embodiment of the present invention, Ubuntu 20.04 LTS runs on the Linux 5.11 kernel, equipped with an Inter(R) Core(TM) 3.20 GHz processor; the design is implemented based on LevelDB, and YCSB is used to generate data sets for experimental testing. To more intuitively evaluate the performance of the present invention, three systems, namely LevelDB, RocksDB, and DiffKV, were used for comparison during the experiment. The MemTable of all systems participating in the experiment was set to 4MB, the SSTable file was set to 2MB, and a 500MB Block Cache was set for all of them. According to the experimental results, under the workload dominated by write operations, the throughput of the present invention has a significant increase compared to other systems, with the throughput being about 1.2 - 1.7 times, 1.1 - 1.5 times, and 1.1 - 1.8 times that of LevelDB, RocksDB, and DiffKV respectively; under the workload dominated by read operations, the performance of the present invention can also be comparable to that of LevelDB, RocksDB, and DiffKV.
[0044] See Figure 5 As shown, it is the system structure diagram of the embodiment of the present invention, including:
[0045] Trigger module 501, a trigger module that triggers a Compaction operation when the amount of data stored in the L layer of the LSM tree reaches its threshold; i Layer stores the data volume reaches its threshold, triggering the Compaction operation;
[0046] Partitioning module 502, selecting the SSTable read tasks of multiple consecutive Compaction operations in the L layer and the L layer and putting them into the task queue; i+1 Layer and the SSTable read tasks of multiple consecutive Compaction operations in the L layer are read into the task queue; i Layer and read into the task queue;
[0047] Compaction module 503, judging whether to perform a large Compaction operation or only modify the metadata according to the key range overlap of the SSTables in the read task queue to implement the request for the SSTables to move between the L layer and the L layer; i Layer and L i+1 Layer to move between requests.
[0048] It can be seen that the present invention proposes a model of early Compaction, which effectively alleviates the additional read-write amplification problem caused by range overlap during continuous Compaction operations. At the same time, a file-level pipelining scheme is added to optimize the proposed early Compaction model, achieving the optimization of the storage system based on the LSM tree. It should be noted here that according to the characteristics of read operations during execution, we perform traditional Compaction operations on the lower-level SSTables, and only adopt the early Compaction method proposed by the present invention at the higher levels.
[0049] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. An early Compaction method applied to the LSM tree structure, characterized in that, Including the following steps: When the amount of data stored in the $i$-th layer $L$ of the LSM tree i reaches its threshold, a Compaction operation is triggered; Select the (i + 1)-th layer L i+1 and L i Read the SSTables of multiple consecutive Compaction operations to be executed subsequently in the layer into the task queue After multiple SSTables are read into the task queue, determine whether to perform a large Compaction operation or only modify the metadata according to different key range overlap situations to implement the request for moving SSTables between the L i layer and the L i+1 layer; Determining whether to perform a large Compaction operation or only modify metadata according to different key range overlap situations to implement the request for moving SSTables between the L i layer and the L i+1 layer, including the following steps: The task queue is divided into two operation queues according to whether there is a key range overlap in the SSTable. If there is no range overlap between the SSTables in layer L i and layer L i+1 , they are assigned to the first operation queue; if there is a range overlap between the SSTables in layer L i and layer L i+1 , they are assigned to the second operation queue; For the first operation queue, only modify the metadata of level L i and level L i+1 , and then insert the SSTable of level L i into the corresponding position of level L i+1 ; for the second operation queue, merge multiple subsequent consecutive Compaction operation tasks with the current Compaction operation task into a large Compaction operation task for execution together, to achieve the migration of SSTable from level L i and level L i+1 . For the second operation queue, in order to accelerate the write-back speed of the SSTable, the write-back of the previous SSTable file and the generation of the next SSTable file are performed simultaneously; The simultaneous write-back of the previous SSTable file and the generation of the next SSTable file are implemented by adopting a file-granularity pipelining scheme, which specifically includes the following steps: Creation step: According to the task information returned in the task queue, an iterator is created for the corresponding SSTable file in the task for subsequent traversal of relevant data, and a file data structure of the SSTable is created in the memory for receiving the subsequent generated ordered data blocks; Writing step: The data is traversed according to the iterator, the expired key-value pair information is deleted, the valid data is retained, and the valid data is organized into multiple data blocks and written into the created SSTable file data structure; When the capacity of the SSTable to store data blocks reaches its threshold, after adding the corresponding metadata to the SSTable, it is written back to the L layer of the disk; during the process of writing the previous SSTable back to the disk, the iterator also continues to traverse the valid data and creates a new SSTable file data structure in memory to receive new valid data blocks; when all the valid data is written back to the L layer in an orderly manner, the operation of early Compaction is completed. i+1 When the capacity of the SSTable to store data blocks reaches its threshold, after adding the corresponding metadata to the SSTable, it is written back to the L layer of the disk; during the process of writing the previous SSTable back to the disk, the iterator also continues to traverse the valid data and creates a new SSTable file data structure in memory to receive new valid data blocks; when all the valid data is written back to the L layer in an orderly manner, the operation of early Compaction is completed. i+1 When all the valid data is written back to the L layer in an orderly manner, the operation of early Compaction is completed.
2. The early Compaction method applied to the LSM tree structure according to claim 1, characterized in that, There are three situations in the second operation queue, and different processing is performed according to different situations, including: For the intersecting and non - intersecting relationships, find the position where the newly generated SSTable after the Compaction operation task is executed is stored in level L i+1 such that it is globally ordered in level L i+1 ; among them, for SSTables that have range overlaps with multiple Compaction operations, they only need to be added to the queue once; For the separated relationship, a stop flag for SSTable interval is set. When generating a new SSTable, the SSTable stops writing before the minimum key of the interval SSTable, and then a new SSTable is created to receive the subsequent data writing.
3. An early Compaction system applied to the LSM tree structure, characterized in that For implementing the method according to any one of claims 1 to 2, including: Trigger module, when the amount of data stored in the i-th layer L of the LSM tree i reaches its threshold, it triggers the Compaction operation; Partition module, select layer L at the (i + 1)-th level i+1 with L i read the SSTables of multiple consecutive Compaction operations to be executed subsequently in layer L into the task queue After the Compaction module reads multiple SSTables into the task queue, it determines whether to perform a large Compaction operation or only modify the metadata according to different key range overlap situations to implement the request for SSTable to move between L i layer and L i+1 layer.
Citation Information
Patent Citations
Storage structure of LSM tree based on NVM and data storage method thereof
CN113821177A
Multistage cooperative compression method and related device for satellite-borne hybrid key value storage system
CN117453141A