LSM tree-based compaction device, method, and computer program

By employing the Simple Copy function of ZNS SSDs for simple copy compaction on overlapping data blocks, the method addresses the write stall issue and enhances compaction efficiency in LSM trees, leading to improved performance and reduced disk overhead.

WO2025135673A1PCT designated stage expired Publication Date: 2025-06-26RES & BUSINESS FOUND SUNGKYUNKWAN UNIV +1
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/020205
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-19
Filing Date
2024-12-10
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

The compaction process in LSM trees often leads to write stalls due to insufficient compaction throughput, resulting in unpredictable performance degradation and increased risk of out-of-memory errors.

Method used

The proposed solution involves a method for accelerating compaction by using the Simple Copy function of ZNS SSDs to perform simple copy compaction on overlapping data blocks, reducing unnecessary data movement between disk and memory.

Benefits of technology

This approach reduces disk read/write overhead, alleviates write stalls, and improves overall performance by maintaining predictable performance and increasing throughput during the compaction process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024020205_26062025_PF_FP_ABST
    Figure KR2024020205_26062025_PF_FP_ABST
Patent Text Reader

Abstract

The present invention relates to a technology for accelerating the compaction of an LSM tree and alleviating a write stall problem by utilizing a simple copy function of an NS. According to one aspect of the present invention, the LSM tree-based compaction method comprises the steps of: on the basis of key values of K-level SSTable including a plurality of data blocks and (K+1)-level SSTable including a plurality of data blocks, searching for one or more overlapping data block groups including a plurality of data blocks having key values in an overlapping relationship; calculating overlapping values between the plurality of data blocks included in each of the overlapping data block groups; and controlling to perform simple copy compaction on the plurality of data blocks included in each of the overlapping data block groups when the overlapping values are smaller than a preset reference value.
Need to check novelty before this filing date? Find Prior Art

Description

LSM tree-based compaction device, method and computer program

[0001] The present invention relates to a technology for accelerating the compaction of an LSM tree and alleviating the write stall problem by utilizing the Simple Copy function of a ZNS SSD.

[0002] The LSM tree is a data structure optimized for write-intensive workloads and is used in various NoSQL platforms such as Cassandra, HBase, RocksDB, and MongoDB. When a client requests to write a key-value pair, the LMS inserts the key-value pair into a MemTable in memory. The MemTable uses data structures such as a skip list to keep the key-value pairs sorted. When the MemTable reaches a certain size, the key-value pairs in the MemTable are converted into an array and stored as a file on disk. This file is called an SSTable, and this process is called flushing.

[0003] An SSTable is largely composed of data blocks, filter blocks, and index blocks. Data blocks store key-value pairs in sorted order. The filter block stores a Bloom filter, a probabilistic data structure used to determine whether a key belongs to each data block. The index block stores the starting address and key value of each data block.

[0004] LSM hierarchically manages SSTables by dividing them into several levels from Level 0 to Level n. Each level contains multiple SSTables, with levels closer to Level 0 being called lower levels and levels closer to Level n being called higher levels. Each level has a limited maximum size, and the size generally increases as it goes up. SSTables that are flushed from memory are stored in Level 0 and are gradually diluted to higher levels through a background job called Compaction. Compaction refers to the process of selecting victim SSTables from Level k when the size of a specific Level k exceeds a set threshold, and then merging them with all Level k+1 SSTables whose key ranges overlap with the victim SSTable to create new Level k+1 SSTables.

[0005] Compaction is an essential operation that allows LSM-trees to hierarchically organize key-value pairs on disk, providing high write performance while ensuring adequate read performance. However, if compaction throughput cannot keep up with user insert throughput, new write requests will continue to pile up in memory, which can lead to out-of-memory errors. Furthermore, if compaction is slow, multiple unsorted SSTables may exist, which can degrade search performance. Therefore, most LSM-tree-based key-value stores, such as HBase, LevelDB, RocksDB, and Cassandra, throttle or block client writes when the number of SSTables at each level exceeds a certain threshold until the compaction thread secures space. This is called a write stall. Write stall causes unpredictable performance degradation because it arbitrarily limits the client write rate, which also degrades overall performance.

[0006] Meanwhile, SSDs are widely used as secondary storage due to their faster processing speeds than HDDs. However, due to the different write and erase units of NAND Flash and the inability to overwrite, Garbage Collection (GC) is inevitable, which has been pointed out as a problem that lowers the quality of service. To solve this problem, the ZNS SSD standard was defined. ZNS divides the SSD space into zones of a certain size and enforces sequential writing for each zone, thereby minimizing the role of the FTL and eliminating GC. Some ZNSs provide a function called Simple Copy, which offloads data copying to the controller to assist the user's GC, enabling data copying from disk to disk without moving data to host memory.

[0007] Prior art literature

[0008] (Public Patent Publication) No. 2023-0096180, Spatial LSM tree device and method for blockchain-based geospatial point data indexing.

[0009] An object of the present invention is to provide a technology for accelerating compaction and alleviating the write stall problem by reducing disk read / write occurring during the compaction process of LSM.

[0010] In addition, it is an object of the present invention to improve a technology for saving disk bandwidth by reducing data movement between disk and memory.

[0011] According to one aspect of the present invention, a compaction method based on an LSM tree may include: a step of searching for at least one overlapping data block group including a plurality of data blocks having overlapping key values ​​based on a K-level SSTable including a plurality of data blocks and a K+1-level SSTable including a plurality of data blocks; a step of calculating an overlap value between the plurality of data blocks included in each of the overlapping data block groups; and a step of controlling to perform a simple copy compaction on the plurality of data blocks included in each of the overlapping data block groups if the overlap value is less than a preset reference value.

[0012] In one embodiment, the plurality of data blocks included in each overlapping data block group may be composed of at least one data block of a K-level SSTable and at least one data block of a K+1-level SSTable that contain the same key value.

[0013] In one embodiment, the step of calculating a duplicate value may calculate the duplicate value by comparing Bloom filters of a plurality of data blocks included in each overlapping data block group.

[0014] In one embodiment, the step of calculating a duplicate value may calculate the duplicate value based on a Hamming distance between Bloom filters of a plurality of data blocks included in each overlapping data block group.

[0015] In one embodiment, the step of calculating a duplicate value may calculate the number of common bits set to 1 in a Bloom filter of a plurality of data blocks included in each overlapping data block group as the duplicate value.

[0016] In one embodiment, the method may further include a step of controlling to perform a merge sort using a buffer memory for a plurality of data blocks included in each overlapping data block group if the duplicate value is greater than or equal to a preset reference value.

[0017] In one embodiment, a simple copy compaction may be a compaction using a simple copy command according to a Zoned Namespace (ZNS) SSD.

[0018] In one embodiment, the step of searching for at least one group of overlapping data blocks further searches for non-overlapping data blocks with non-duplicate key values, and the step of controlling to perform a simple copy compaction can control the simple copy compaction for the non-overlapping data blocks.

[0019] In addition, according to another aspect of the present invention, an LSM tree-based compaction device may include: an overlapping data block search unit that searches for at least one overlapping data block group including a plurality of data blocks having overlapping key values ​​based on a K-level SSTable including a plurality of data blocks and a K+1-level SSTable including a plurality of data blocks; a data block overlap value calculation unit that calculates an overlap value between the plurality of data blocks included in each of the overlapping data block groups; and a compaction control unit that controls to perform a simple copy compaction on the plurality of data blocks included in each of the overlapping data block groups when the overlap value is less than a preset reference value.

[0020] According to one aspect of the present invention, it is possible to accelerate compaction and alleviate the write stall problem by reducing disk read / write occurring during the compaction process of LSM.

[0021] Additionally, according to another aspect of the present invention, it is possible to save disk bandwidth by reducing data movement between disk and memory.

[0022] Thus, the present invention enables predictable performance of LSM and increased overall performance.

[0023] Figure 1 is a diagram for explaining conventional LSM tree-based compaction.

[0024] FIG. 2 is a block diagram of an LSM tree-based compaction device according to one embodiment of the present invention.

[0025] FIG. 3 is a flowchart of an LSM tree-based compaction method according to one embodiment of the present invention.

[0026] FIG. 4 is a diagram for explaining compaction based on an LSM tree according to one embodiment of the present invention.

[0027] FIG. 5 is a block diagram of an LSM tree-based compaction device according to another embodiment of the present invention.

[0028] The advantages and features of the present invention, and the methods for achieving them, will become clearer with reference to the embodiments described in detail below with the accompanying drawings. However, the present invention is not limited to the embodiments disclosed below and may be implemented in various forms. These embodiments are provided solely to ensure that the disclosure of the present invention is complete and to fully inform those skilled in the art of the scope of the invention, and the scope of the present invention is defined solely by the claims.

[0029] In describing embodiments of the present invention, specific descriptions of known functions or configurations will be omitted unless actually necessary. Furthermore, the terms described below are defined based on their functions in the embodiments of the present invention and may vary depending on the intent or custom of the user or operator. Therefore, their definitions should be based on the overall content of this specification.

[0030] The terms ‘…bu’, ‘…gi’, etc. used hereinafter mean a unit that processes at least one function or operation, which can be implemented by hardware, software, or a combination of hardware and software.

[0031]

[0032] Figure 1 is a diagram for explaining conventional LSM tree-based compaction.

[0033] Referring to Figure 1, it shows the problem of the existing compaction process in the LSM Tree. SSTable A of Level K, which was selected as the victim, overlaps the key ranges of SSTables B, C, and D of the next levels. Let us assume that the smallest key stored in data block A1 is smaller than the largest key stored in data block B4, and the largest key stored in data block A4 is smaller than the largest key stored in data block D1. In other words, the data stored in the blocks marked in gray indicate that their key values ​​overlap with each other. In other words, the key ranges of the data stored in data blocks B1, B2, B3, D2, D3, and D4 do not overlap with those of other blocks. These blocks are called non-overlapping blocks.

[0034] Existing compaction considers key range overlap at the SSTable level, not the block level, and performs a merge sort if key ranges overlap. Therefore, data stored in data blocks B1, B2, B3, D2, D3, and D4 are read from ZNS storage, copied to DRAM, and then stored in blocks E1, E2, E3, H2, H3, and H4 of ZNS storage, even though their key ranges do not overlap with those of other blocks and thus do not require merge sorting. This unnecessarily incurs the overhead of copying data from disk to memory.

[0035] ZNS's Simple Copy command allows blocks in ZNS storage to be copied to blocks in another ZNS without having to copy them to DRAM. This patent proposes a method to avoid unnecessary memory copies by using the Simple Copy command to directly copy blocks B1, B2, B3, D2, D3, and D4 to blocks E1, E2, E3, H2, H3, and H4.

[0036] Meanwhile, if the key distribution is not dense (the keys are widely spaced), data blocks in the middle of the key range can also be written to the newly created SSTable without any changes. For example, suppose the key range of data block C3 in the above example is 100-110. If SSTable A does not have a key between 100 and 110, the newly created SSTable will have data block G1 with the same contents as C3. Since there are no data blocks with overlapping ranges, non-overlapping data blocks whose contents remain unchanged even after compaction can exist anywhere in the SSTable.

[0037] Therefore, the present invention aims to provide a technology for accelerating compaction by eliminating the overhead of reading such non-overlapping data blocks into host memory, calculating them, and then writing them back to the SSD.

[0038] In addition, the present invention seeks to provide a technology that can be applied to overlapping data blocks in which key ranges overlap.

[0039]

[0040] FIG. 2 is a block diagram of an LSM tree-based compaction device according to one embodiment of the present invention.

[0041] Referring to FIG. 2, an LSM tree-based compaction device (2000) according to one embodiment of the present invention may include an overlapping data block search unit (2100), a data block duplication value calculation unit (2200), and a compaction control unit (2300).

[0042] The overlapping data block search unit (2100) can search for data blocks included in an SSTable by checking the key values ​​included in the SSTable.

[0043] In one embodiment, the overlapping data block search unit (2100) can search for at least one overlapping data block group including a plurality of data blocks having overlapping key values ​​based on key values ​​of a K-level SSTable including a plurality of data blocks and a K+1-level SSTable including a plurality of data blocks.

[0044] In one embodiment, the plurality of data blocks included in each overlapping data block group may be composed of at least one data block of a K-level SSTable and at least one data block of a K+1-level SSTable that contain the same key value.

[0045] In one embodiment, the overlapping data block search unit (2100) can determine, based on the search result for duplication of data blocks, whether the data blocks included in the K-level SSTable and the K+1-level SSTable are overlapping data block groups with overlapping key values ​​or non-overlapping data blocks with non-overlapping key values.

[0046] The data block duplication value calculation unit (2200) can calculate duplication values ​​for data blocks included in an overlapping data block group.

[0047] In one embodiment, the data block duplication value calculation unit (2200) can calculate a duplication value by comparing Bloom filters of multiple data blocks included in each overlapping data block group.

[0048] In one embodiment, the data block duplication value calculation unit (2200) can calculate the duplication value based on the Hamming distance between Bloom filters of a plurality of data blocks included in each overlapping data block group.

[0049] In one embodiment, the data block duplication value calculation unit (2200) can calculate the number of common bits set to 1 in a Bloom filter of a plurality of data blocks included in each overlapping data block group as a duplication value.

[0050] The compaction control unit (2300) can determine whether the duplication value for data blocks included in the overlapping data block group is less than a preset reference value.

[0051] In one embodiment, the preset reference value may be preset by the size of the storage space, the user, etc.

[0052] In addition, the compaction control unit (2300) can control simple copy compaction to be performed on data blocks included in an overlapping data block group or non-overlapping data blocks whose duplication value is less than a preset reference value.

[0053] In one embodiment, the compaction control unit (2300) can control to perform simple copy compaction on a storage device such as an SSD.

[0054] In one embodiment, a simple copy compaction may be a compaction using a simple copy command according to a Zoned Namespace (ZNS) SSD.

[0055] In addition, the compaction control unit (2300) can control to perform a merge sort using a buffer memory for a plurality of data blocks included in each overlapping data block group if the duplicate value is greater than or equal to a preset reference value.

[0056] In one embodiment, the compaction control unit (2300) may control a storage device such as an SSD to perform a merge sort.

[0057]

[0058] FIG. 3 is a flowchart of an LSM tree-based compaction method according to one embodiment of the present invention.

[0059] The method described below is described as an example performed by an LSM tree-based compaction device (2000) illustrated in FIG. 2.

[0060] In step S3100, the LSM tree-based compaction device (2000) can search for data blocks included in the SSTable by checking the key values ​​included in the SSTable.

[0061] In one embodiment, an LSM tree-based compaction device (2000) can search for at least one overlapping data block group including a plurality of data blocks having overlapping key values ​​based on a K-level SSTable including a plurality of data blocks and a K+1-level SSTable including a plurality of data blocks.

[0062] In one embodiment, the plurality of data blocks included in each overlapping data block group may be composed of at least one data block of a K-level SSTable and at least one data block of a K+1-level SSTable that contain the same key value.

[0063] In step S3200, the LSM tree-based compaction device (2000) can determine, based on the result of the search for duplication of data blocks, whether the data blocks included in the K-level SSTable and the K+1-level SSTable are overlapping data block groups with overlapping key values ​​or non-overlapping data blocks with non-overlapping key values.

[0064] In one embodiment, the LSM tree-based compaction device (2000) may perform step S3500 for non-overlapping data blocks and step S3300 for overlapping data block groups.

[0065] In step S3300, the LSM tree-based compaction device (2000) can calculate a redundancy value for data blocks included in an overlapping data block group.

[0066] In one embodiment, the LSM tree-based compaction device (2000) can calculate a redundancy value by comparing Bloom filters of multiple data blocks included in each overlapping data block group.

[0067] In one embodiment, the LSM tree-based compaction device (2000) can calculate a redundancy value based on a Hamming distance between Bloom filters of a plurality of data blocks included in each overlapping data block group.

[0068] In one embodiment, the LSM tree-based compaction device (2000) can calculate the number of common bits set to 1 in a Bloom filter of a plurality of data blocks included in each overlapping data block group as a redundancy value.

[0069] In step S3400, the LSM tree-based compaction device (2000) can determine whether the duplication value for data blocks included in the overlapping data block group is less than a preset reference value. If the duplication value is less than the preset reference value, the LSM tree-based compaction device (2000) can perform step S3500, and if the duplication value is greater than or equal to the preset reference value, the device can perform step S3600.

[0070] In one embodiment, the preset reference value may be preset by the size of the storage space, the user, etc.

[0071] In step S3500, the LSM tree-based compaction device (2000) can control simple copy compaction to be performed on data blocks included in an overlapping data block group or non-overlapping data blocks whose duplicate value is less than a preset reference value.

[0072] In one embodiment, an LSM tree-based compaction device (2000) can be controlled to perform simple copy compaction on a storage device such as an SSD.

[0073] In one embodiment, a simple copy compaction may be a compaction using a simple copy command according to a Zoned Namespace (ZNS) SSD.

[0074] In step S3600, the LSM tree-based compaction device (2000) can be controlled to perform a merge sort using a buffer memory for a plurality of data blocks included in each overlapping data block group if the duplicate value is greater than or equal to a preset reference value.

[0075] In one embodiment, an LSM tree-based compaction device (2000) can be controlled to perform merge sort on a storage device such as an SSD.

[0076]

[0077] FIG. 4 is a diagram for explaining compaction based on an LSM tree according to one embodiment of the present invention.

[0078] Referring to FIG. 4, an example of Simple Copy Compaction according to one embodiment of the present invention is shown. When SSTable A and SSTable B are compacted, the LSM tree-based compaction device (2000) knows the minimum and maximum keys of each data block through the index block, so it can confirm that blocks B1, A2, B3, and A4 are non-overlapping data blocks, and thus can perform a simple copy.

[0079] Meanwhile, blocks A3 and B4 overlap in scope, but only to a very small degree. Therefore, even if merge sort is not performed on these two blocks, search performance may not be significantly affected. The present invention utilizes this opportunity to perform a simple copy on overlapping data blocks, provided the degree of overlap is not significant.

[0080] The LSM tree-based compaction device (2000) determines whether to perform a simple copy when the degree of overlap is not large by comparing the Bloom filters of overlapping data blocks. At this time, the overlap value of the two filters is calculated, and the similarity of the two filters can be determined by using the Hamming distance or the number of common bits set to 1 in the two filters. If the two filters are significantly different, i.e., the overlap value is small, a simple copy can be performed.

[0081] When performing a simple copy on overlapping data blocks, the ranges of the two data blocks overlap. Finding a key within that range requires accessing both blocks, which can degrade read performance. However, if the overlap between filters is small, a Bloom filter can effectively filter out unnecessary read operations.

[0082] For example, even if the key ranges of data blocks B3 and A4 overlap, if the Bloom filter indicates that the key does not exist in either data block, the query does not need to read both data blocks. In this case, the read operation can skip accessing the corresponding data block. Alternatively, the Bloom filter may indicate that the key is likely to exist only in data block A4. In this case, data block B3 does not need to be read. Determining whether to perform a simple copy based on duplicate values ​​can reduce the number of cases in which both data blocks must be accessed when searching for a key.

[0083]

[0084] FIG. 5 is a block diagram of an LSM tree-based compaction device according to another embodiment of the present invention.

[0085] As illustrated in FIG. 5, the LSM tree-based compaction device (2000) may include at least one of a processor (5100), a memory (5200), a storage (5300), a user interface input unit (5400), and a user interface output unit (5500), which may communicate with each other via a bus (5600). In addition, the LSM tree-based compaction device (2000) may also include a network interface (5700) for connecting to a network. The processor (5100) may be a CPU or a semiconductor device that executes processing instructions stored in the memory (5200) and / or the storage (5300). The memory (5200) and the storage (5300) may include various types of volatile / non-volatile storage media. For example, the memory may include a ROM (5240) and a RAM (5250).

[0086]

[0087] The devices described above may be implemented as hardware components, software components, and / or a combination of hardware components and software components. For example, the devices and components described in the embodiments may be implemented using one or more general-purpose computers or special-purpose computers, such as, for example, a processor, a controller, an arithmetic logic unit (ALU), a digital signal processor, a microcomputer, a field programmable array (FPA), a programmable logic unit (PLU), a microprocessor, or any other device capable of executing instructions and responding. The processing device may execute an operating system (OS) and one or more software applications running on the operating system.

[0088] Additionally, the processing device may access, store, manipulate, process, and generate data in response to the execution of software. For ease of understanding, the processing device is sometimes described as being used alone; however, those skilled in the art will appreciate that the processing device may include multiple processing elements and / or multiple types of processing elements. For example, the processing device may include multiple processors, or a processor and a controller. Other processing configurations, such as parallel processors, are also possible.

[0089] Software may include a computer program, code, instructions, or a combination of one or more of these, which may configure a processing device to perform a desired operation or may, independently or collectively, command the processing device. The software and / or data may be permanently or temporarily embodied in any type of machine, component, physical device, virtual equipment, computer storage medium or device, or transmitted signal wave, for interpretation by the processing device or for providing instructions or data to the processing device. The software may also be distributed over networked computer systems and stored or executed in a distributed manner. The software and data may be stored on one or more computer-readable recording media.

[0090]

[0091] The above description is merely an illustrative example of the technical idea of ​​the present invention, and those skilled in the art will appreciate that various modifications and variations can be made without departing from the essential quality of the present invention. Therefore, the embodiments disclosed in this specification are intended to illustrate rather than limit the technical idea of ​​the present invention, and the scope of the technical idea of ​​the present invention is not limited by these embodiments. The scope of protection of the present invention should be interpreted by the following claims, and all technical ideas within a scope equivalent thereto should be interpreted as being included in the scope of the rights of the present invention.

[0092] 2000: LSM tree-based compaction device

[0093] 2100: Overlapping Data Block Search Unit

[0094] 2200: Data block duplicate value calculation unit

[0095] 2300: Compaction Control Unit

Claims

1. A step of searching for at least one overlapping data block group including a plurality of data blocks having overlapping key values ​​based on a K-level SSTable including a plurality of data blocks and a K+1-level SSTable including a plurality of data blocks; A step of calculating a duplicate value between a plurality of data blocks included in each of the above overlapping data block groups; and A step of controlling to perform a simple copy compaction for a plurality of data blocks included in each of the overlapping data block groups if the above duplicate value is less than a preset reference value; A compaction method based on LSM tree, including.

2. In paragraph 1, A compaction method based on an LSM tree, wherein each of the plurality of data blocks included in each of the above overlapping data block groups is composed of at least one data block of the K-level SSTable including the same key value and at least one data block of the K+1-level SSTable.

3. In paragraph 1, The step of calculating the above duplicate values ​​is: An LSM tree-based compaction method for calculating the overlap value by comparing bloom filters of multiple data blocks included in each of the above overlapping data block groups.

4. In paragraph 3, The step of calculating the above duplicate values ​​is: An LSM tree-based compaction method for calculating the overlap value based on the Hamming distance between bloom filters of multiple data blocks included in each of the above overlapping data block groups.

5. In paragraph 4, The step of calculating the above duplicate values ​​is: An LSM tree-based compaction method, which calculates the number of common bits set to 1 in a bloom filter of multiple data blocks included in each of the above overlapping data block groups as the redundancy value.

6. In paragraph 1, A compaction method based on an LSM tree, further comprising a step of controlling to perform a merge sort using a buffer memory for a plurality of data blocks included in each of the overlapping data block groups if the above duplicate value is greater than or equal to a preset reference value.

7. In paragraph 1, The above simple copy compaction is a compaction method based on an LSM tree that uses a simple copy command according to a ZNS (Zoned Namespace) SSD.

8. In paragraph 1, The step of searching for at least one overlapping data block group comprises: Find more non-overlapping data blocks with non-duplicate key values, The steps to control the execution of a simple copy compaction are: An LSM tree-based compaction method for controlling the simple copy compaction for the above non-overlapping data blocks.

9. An overlapping data block search unit that searches for at least one overlapping data block group including a plurality of data blocks having overlapping key values ​​based on a K-level SSTable including a plurality of data blocks and a K+1-level SSTable including a plurality of data blocks; A data block overlap value calculating unit for calculating overlap values ​​between multiple data blocks included in each of the above overlapping data block groups; and A compaction control unit that controls to perform a simple copy compaction on a plurality of data blocks included in each of the overlapping data block groups if the above duplicate value is less than a preset reference value; A compaction device based on an LSM tree, comprising:

10. In paragraph 9, A compaction device based on an LSM tree, wherein each of the plurality of data blocks included in each of the above overlapping data block groups is composed of at least one data block of the K-level SSTable including the same key value and at least one data block of the K+1-level SSTable.

11. In paragraph 9, The above data block duplicate value calculation section is, An LSM tree-based compaction device that calculates the redundancy value by comparing bloom filters of a plurality of data blocks included in each of the above overlapping data block groups.

12. In paragraph 11, The above data block duplicate value calculation section is, An LSM tree-based compaction device that calculates the redundancy value based on a Hamming distance between bloom filters of a plurality of data blocks included in each of the above overlapping data block groups.

13. In paragraph 12, The above data block duplicate value calculation section is, An LSM tree-based compaction device that calculates the number of common bits set to 1 in a bloom filter of a plurality of data blocks included in each of the above overlapping data block groups as the redundancy value.

14. In paragraph 9, An LSM tree-based compaction device further comprising a step of controlling to perform a merge sort using a buffer memory for a plurality of data blocks included in each of the overlapping data block groups if the above duplicate value is greater than or equal to a preset reference value.

15. In paragraph 9, The above simple copy compaction is a compaction device based on an LSM tree that uses a simple copy command according to a ZNS (Zoned Namespace) SSD.

16. In paragraph 1, The above overlapping data block search unit, Find more non-overlapping data blocks with non-duplicate key values, The above compaction control unit, An LSM tree-based compaction device that performs the simple copy compaction on the above non-overlapping data blocks.

17. Memory containing instructions for executing a compaction method based on an LSM tree; and By executing the above command, a step of searching for at least one overlapping data block group including a plurality of data blocks having overlapping key values ​​based on key values ​​of a K-level SSTable including a plurality of data blocks and a K+1-level SSTable including a plurality of data blocks; A step of calculating a duplicate value between a plurality of data blocks included in each of the above overlapping data block groups; and A step of controlling to perform a simple copy compaction for a plurality of data blocks included in each of the overlapping data block groups if the above duplicate value is less than a preset reference value; A processor performing a compaction method based on an LSM tree, comprising: A compaction device based on an LSM tree, comprising:

18. A computer-readable recording medium storing a computer program, The above computer program, when executed by a processor, performs the steps of: searching for at least one overlapping data block group including a plurality of data blocks having overlapping key values ​​based on key values ​​of a K-level SSTable including a plurality of data blocks and a K+1-level SSTable including a plurality of data blocks by executing the instructions; A step of calculating a duplicate value between a plurality of data blocks included in each of the above overlapping data block groups; and A step of controlling to perform a simple copy compaction for a plurality of data blocks included in each of the overlapping data block groups if the above duplicate value is less than a preset reference value; A method for compacting a LSM tree based on a processor, comprising instructions for causing the processor to perform the compaction method. Computer readable recording medium.

19. A computer program stored on a computer-readable recording medium, The above computer program, when executed by a processor, performs the steps of: searching for at least one overlapping data block group including a plurality of data blocks having overlapping key values ​​based on key values ​​of a K-level SSTable including a plurality of data blocks and a K+1-level SSTable including a plurality of data blocks by executing the instructions; A step of calculating a duplicate value between a plurality of data blocks included in each of the above overlapping data block groups; and A step of controlling to perform a simple copy compaction for a plurality of data blocks included in each of the overlapping data block groups if the above duplicate value is less than a preset reference value; A method for compacting a LSM tree based on a processor, comprising instructions for causing the processor to perform the compaction method. Computer program.

Citation Information

Patent Citations

  • Woopung soundproofing pad

    KR1020240124136A

  • Malfunction diagnosis method and apparatus of transmission using supervised learning based artificial intelligence

    KR1020240125801A

  • Memory sytem and operating method thereof

    KR102512571B1

  • Scalable I / O operations on a log-structured merge (LSM) tree

    US20220156231A1