A compression and merging method and system for key-value storage
By distinguishing the upper and the lowest level compression mechanism in the LSM tree, introducing active ordered string tables and merging and compression, the problem of difficult to weigh write overhead, read overhead and storage space overhead in the prior art is solved, and the scalability of the key-value storage system is improved.
Patent Information
- Application Number
- CN202310318032.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-22
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2043-03-22
AI Technical Summary
The existing LSM tree compression strategy is difficult to make a reasonable trade-off between write overhead, read overhead and storage space overhead, resulting in a lack of scalability in key-value storage systems.
Two parameters are introduced to determine the active ordered string table in the upper and lowermost layer respectively. Through merging and compression, all ordered string tables in the current layer are merged with the active ordered string table in the next layer, expanding the design space of the LSM tree and adapting to the changing data load.
By distinguishing between the upper and the lowest level, the overhead and space amplification of target key query, short range search, and long range search are reduced, while maintaining a low update overhead, improving the scalability of the key-value storage system.
Smart Images

Figure CN116340276B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of databases, and particularly to a method and system for compressing and merging key-value storage. Background Art
[0002] A key-value storage system, also known as a key-value database, uses simple key-value pairs as the underlying data model. Key-value storage is generally used as an independent NoSQL system. Different from relational databases, key-value storage does not have a relational query language, but completes data management through simple operations such as inserting, querying, and deleting key-value pairs. In data-intensive applications, key-value storage can provide better write performance and relatively fast query performance. Key-value storage allows horizontal scaling, which cannot be achieved by other types of databases and is highly partitionable.
[0003] Commonly, hash tables, B+ trees, and LSM trees (Log-Structure-Merge Trees) are used as the underlying data structures of key-value systems. An LSM tree is a hierarchical and ordered data storage method based on a hard disk. Here, the hierarchy refers to the storage of memory and hard disk. The common architecture of a key-value storage system based on an LSM tree is as Figure 2 shown.
[0004] To avoid wasting disk space, the compression and merging operation in a key-value storage system is very important. Read amplification, write amplification, and space amplification are three important concepts related to compression. Read amplification refers to the situation where the actual amount of data read during data reading is greater than the real amount of data. Write amplification refers to the situation where the actual amount of data written during data writing is greater than the real amount of data. Space amplification refers to the situation where the disk space actually occupied by data is more than the real size of the data. Compression strategies need to make a trade-off among the three negative effects to adapt to specific application scenarios.
[0005] Existing LSM tree compression strategies are difficult to make a reasonable trade-off among write overhead, read overhead, and storage space overhead, and always make a certain overhead too large, resulting in the lack of scalability of the key-value storage system.
[0006] Therefore, the existing technology still needs to be improved and developed. Summary of the Invention
[0007] The technical problem to be solved by the present invention is to provide a method and system for compressing and merging key-value storage in view of the above-mentioned defects of the existing technology, aiming to solve the problem that existing LSM tree compression strategies are difficult to make a reasonable trade-off among write overhead, read overhead, and storage space overhead, resulting in the lack of scalability of the key-value storage system.
[0008] The technical solution adopted by the present invention to solve the problem is as follows:
[0009] In a first aspect, an embodiment of the present invention provides a compression and merging method for a key-value storage system. The method is applied to a key-value storage system constructed by an LSM tree. The key-value storage system includes several layers for storing key-value pairs. Among them, each non-bottom layer includes a first number of sorted string tables, and the bottom layer includes a second number of the sorted string tables. The method includes:
[0010] Determine a first target layer to be compressed in the key-value storage system, and determine a second target layer located below the first target layer, where the first target layer is not the bottom layer;
[0011] When the second target layer is not the bottom layer, determine whether there is an active sorted string table in the second target layer according to the first number and the total file size of the first target layer;
[0012] When the second target layer is the bottom layer, determine whether there is the active sorted string table in the second target layer according to the second number and the total file size of the first target layer;
[0013] When there is one, after sorting and merging all the sorted string tables in the first target layer, then sort and merge them with the active sorted string table to obtain a first merged sorted string table; move the first merged sorted string table to the second target layer;
[0014] When there is none, after sorting and merging all the sorted string tables in the first target layer, obtain a second merged sorted string table; move the second merged sorted string table to the second target layer.
[0015] In an implementation manner, the method for determining the first target layer includes:
[0016] For each non-bottom layer in the key-value storage system, when writing the last sorted string table, use this layer as the first target layer.
[0017] In an implementation manner, determining whether there is an active sorted string table in the second target layer according to the first number and the total file size of the first target layer includes:
[0018] Define a first threshold as the product of the first number and the total file size of the first target layer;
[0019] When there is a sorted string table in the second target layer whose file size is greater than or equal to the first threshold, determine that there is the active sorted string table in the second target layer, and determine the active sorted string table from the sorted string tables whose file size is greater than or equal to the first threshold.
[0020] In one embodiment, the method further includes:
[0021] When there is no ordered string table in the second target layer whose file size is greater than or equal to the first threshold, it is determined that there is no active ordered string table in the second target layer.
[0022] In one embodiment, the first quantity is the difference between the ratio of the total file size of the second target layer and the first target layer and 1.
[0023] In one embodiment, determining whether there is an active ordered string table in the second target layer according to the second quantity and the total file size of the first target layer includes:
[0024] Defining a second threshold as the product of the second quantity and the total file size of the first target layer;
[0025] When there is an ordered string table in the second target layer whose file size is greater than or equal to the second threshold, it is determined that there is an active ordered string table in the second target layer, and the active ordered string table is determined from the ordered string tables whose file size is greater than or equal to the second threshold.
[0026] In one embodiment, the method further includes:
[0027] When there is no ordered string table in the second target layer whose file size is greater than or equal to the second threshold, it is determined that there is no active ordered string table in the second target layer.
[0028] In one embodiment, the second quantity is 1.
[0029] In a second aspect, an embodiment of the present invention further provides a key-value storage system that can be compressed and merged. Among them, the key-value storage system is constructed based on the LSM tree, and the key-value storage system includes:
[0030] Several layers for storing key-value pairs. Among them, each non-bottom layer includes a first quantity of ordered string tables, and the bottom layer includes a second quantity of the ordered string tables;
[0031] A compression control module for determining a first target layer to be compressed in the key-value storage system and determining a second target layer located below the first target layer, where the first target layer is not the bottom layer;
[0032] When the second target layer is not the bottom layer, it is determined whether there is an active ordered string table in the second target layer according to the first quantity and the total file size of the first target layer;
[0033] When the second target layer is the bottom layer, it is determined whether the active ordered string table exists in the second target layer according to the second quantity and the total file size of the first target layer;
[0034] When it exists, after all the ordered string tables of the first target layer are merged in an orderly manner, they are then merged with the active ordered string table in an orderly manner to obtain a first merged ordered string table; the first merged ordered string table is moved into the second target layer;
[0035] When it does not exist, all the ordered string tables of the first target layer are merged in an orderly manner to obtain a second merged ordered string table; the second merged ordered string table is moved into the second target layer.
[0036] In a third aspect, an embodiment of the present invention further provides a computer-readable storage medium, on which multiple instructions are stored, where the instructions are suitable for being loaded and executed by a processor to implement the steps of the compression and merging method of the key-value storage system described in any one of the above.
[0037] Advantages of the present invention: The embodiments of the present invention distinguish the compression mechanisms of the upper layer and the bottom layer, and introduce two parameters to respectively determine the active ordered string tables in the upper layer and the bottom layer. During merge compression, all the ordered string tables of the current layer and the active ordered string tables of the next layer are merged and compressed and then flushed to the next layer. By changing the two parameters, it is possible to switch between different merge compression strategies, expanding the design space of the LSM tree, and thus better adapting to changing data loads. It solves the problem that it is difficult to make a reasonable trade-off among the write overhead, read overhead, and storage space overhead in the LSM tree compression strategy in the prior art, resulting in the lack of scalability of the key-value storage system. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0039] Figure 1 It is a schematic flowchart of the compression and merging method of the key-value storage system provided by the embodiment of the present invention.
[0040] Figure 2 It is a common architecture diagram of the key-value storage system based on the LSM tree provided by the embodiment of the present invention.
[0041] Figure 3It is a schematic diagram showing the trade - off between two mainstream compression strategies of the existing LSM tree provided by the embodiments of the present invention.
[0042] Figure 4 It is a schematic diagram of the lazy hierarchical compression strategy proposed by the present invention provided by the embodiments of the present invention.
[0043] Figure 5 It is a schematic diagram of implementing the lazy hierarchical compression strategy in the LSM tree provided by the embodiments of the present invention.
[0044] Figure 6 It is a schematic diagram of the modules of a key - value storage system with compressible merging provided by the embodiments of the present invention.
[0045] Figure 7 It is a principle block diagram of a terminal provided by the embodiments of the present invention. Detailed implementation manners
[0046] The present invention discloses a method and system for compression and merging of key - value storage. To make the objectives, technical solutions and effects of the present invention clearer and more definite, the following further describes the present invention in detail with reference to the accompanying drawings and by way of examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0047] Those skilled in the art of the present technology can understand that unless specifically stated, the singular forms "a", "an", "the" and "said" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the specification of the present invention means the presence of the described features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or their groups. It should be understood that when we say an element is "connected" or "coupled" to another element, it can be directly connected or coupled to other elements, or there may also be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The phrase "and / or" used herein includes all or any unit and all combinations of one or more related listed items.
[0048] Those skilled in the art of the present technology can understand that unless otherwise defined, all terms (including technical terms and scientific terms) used herein have the same meaning as the general understanding of those of ordinary skill in the field to which the present invention belongs. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art, and will not be interpreted with an idealized or overly formal meaning unless specifically defined as here.
[0049] Existing LSM tree compression strategies include the Size Tired Compaction Strategy (STCS) and the Leveled Compaction Strategy (LCS). Between STCS and LCS, the most important factor controlling the compression frequency is the ratio T = size(Li+1) / size(Li) of the sizes between two adjacent layers, where size(Li) represents the total file size of layer Li. Under STCS, when the value of T increases, the frequency of compression operations decreases and eventually becomes a logging mode. In this final state, the write amplification effect is minimal, but the query overhead is high and the space amplification problem is the most serious. Under LCS, when the value of T decreases, the compression operations are more frequent and eventually become an ordered array recording mode. In this final state, query operations can be completed in one read operation through a fence pointer and redundant key-value pairs are not read, so the read efficiency is high, but the write amplification problem is serious.
[0050] Figure 3 Shows the trade-off between the two compression strategies, where the X-axis is the update overhead, the Y-axis is the query overhead, the curve is generated by the LSM tree selecting different values of T, and the two endpoints represent the logging and ordered array recordings. Facing diverse application scenarios and complex data loads, the existing key-value storage systems often focus on one aspect while neglecting the other when tuning the LSM tree, and as the data scale grows, it may cause the curve to shift upward and to the right, resulting in poor scalability of the system performance. It can be seen that regardless of whether the STCS or LCS compression strategy is adopted, the frequency of compression operations in the LSM tree controls the trade-off between write overhead, read overhead, and storage space overhead. The existing designs lack a reasonable trade-off, and one of the above-mentioned overheads is always too large, resulting in the lack of scalability of the key-value storage system.
[0051] In view of the above-mentioned defects of the prior art, the present invention provides a compression and merging method for a key-value storage system, as Figure 1 shown, the method is applied to a key-value storage system constructed by an LSM tree, the key-value storage system includes several layers for storing key-value pairs, wherein each non-bottom layer includes a first number of ordered string tables, and the bottom layer includes a second number of the ordered string tables, the method includes:
[0052] Step S100, determining a first target layer to be compressed in the key-value storage system, and determining a second target layer located below the first target layer, wherein the first target layer is not the bottom layer;
[0053] Step S200, when the second target layer is not the bottom layer, judging whether there is an active ordered string table in the second target layer according to the first number and the total file size of the first target layer;
[0054] Step S300: When the second target layer is the bottom layer, determine whether the active ordered string table exists in the second target layer according to the second quantity and the total file size of the first target layer;
[0055] Step S400: When it exists, after orderly merging all the ordered string tables of the first target layer, orderly merge them with the active ordered string table to obtain a first merged ordered string table; move the first merged ordered string table to the second target layer;
[0056] Step S500: When it does not exist, orderly merge all the ordered string tables of the first target layer to obtain a second merged ordered string table; move the second merged ordered string table to the second target layer.
[0057] Specifically, the common architecture of the key-value storage system of the LSM tree is as Figure 2As shown in Figure 1, there are four basic operations in the key-value storage system, namely, point lookup, update, short range lookup, and long range lookup. Short range lookup means that the key-value pair to be queried can be read by reading one file block in each ordered string table (SSTable), while long range lookup requires reading multiple file blocks to complete. The existing LSM tree compression strategy is difficult to make a reasonable trade-off between write overhead, read overhead, and storage space overhead, and always makes one kind of overhead too large. The root of this problem is that the update cost, query cost, and space magnification rate of each operation in different layers are different. Among them, update cost: The update of LSM tree is completed by insertion operation, and the insertion of key-value pairs may cause compression operation. The cost of compression operation increases exponentially with the depth of the layer, and the frequency of compression operation decreases exponentially with the depth of the layer. Therefore, the I / O cost caused by update operation is roughly equal at each layer. Target key query: With the help of Bloom filters and fence pointers, the query of each SSTable can be completed by reading a file block. Since the target key is likely to appear only in the bottom layer, the greater the false positive probability of the upper layer, the greater the I / O overhead of the target key query. Long range query: The capacity of each layer of the LSM tree grows exponentially, and the bottom layer contains most of the data. Therefore, the key-value pairs of long range queries are likely to be distributed in the bottom layer, and most of the I / O is performed in the bottom layer. Short range query: Regardless of the layer, short range queries can be completed by reading a file block. Since the maximum number of SSTables in each layer is fixed and each SSTable only reads one file block, the overhead of short range queries is equal in all layers. Space amplification: When the system adopts the STCS compression strategy and the key-value pairs in the bottom layer are redundant versions of their upper layers, the space amplification problem is most serious. Because the probability of redundant key-value pairs appearing in the bottom layer is the highest, the bottom layer has the greatest impact on the space amplification problem.
[0058] It can be seen that the overhead of target key query and long range query mainly comes from the bottom layer, and the space magnification rate mainly depends on the bottom layer. Therefore, most of the compression operations in the upper layer have little effect on performance improvement. To significantly improve performance, it is necessary to change the compression strategy of the bottom layer. The compression and merging method of the key-value storage system provided in this embodiment is called Lazy Leveling Compaction Strategy (LLCS for short), which is mainly obtained by distinguishing the compression mechanism of the upper layer and the bottom layer. Specifically, Figure 4As shown, in this embodiment, the LSM tree is divided into an upper layer and a bottom layer, and two parameters are introduced to determine the best balance point in the design space to adapt to the mixed and variable data load. These two parameters are respectively the number of SSTables in each layer of the upper layer, that is, the first number (K); the number of SSTables in the bottom layer, that is, the second number (Z). The compression and merging method provided in this embodiment mainly acts on two adjacent layers in the key-value storage system. Assume that the upper layer of the two adjacent layers is the first target layer to be compressed, and the lower layer is the second target layer. An active SSTable is set in each layer in this embodiment. For each layer in the key-value storage system except the bottom layer, the active SSTable in this layer is determined based on the K value and the total file size of its upper layer; while the active SSTable in the bottom layer is determined based on the Z value and the total file size of its upper layer. The compression and merging operation is specifically as follows: After orderly merging all the SSTables in the first target layer, it is then orderly merged with the active SSTable in the second target layer to generate a new SSTable, replacing the SSTable in the lower layer participating in the merging. If there is no SSTable in the second target layer with a file size not lower than the corresponding threshold, that is, there is no active SSTable, then after orderly merging all the SSTables in the first target layer, it is no longer merged with the SSTables in the second target layer, but the merged SSTable is directly added to the second target layer. In this embodiment, by setting different K values and Z values, it is possible to switch between different merging and compression strategies, expanding the design space of the LSM tree. When the K value and the Z value decrease, the target key query overhead, short-range search overhead, long-range search overhead, and space amplification factor will gradually decrease, and the overall update overhead will gradually approach the update overhead of the LCS. When the K value and the Z value increase, the overall update overhead will gradually decrease, but the target key query overhead, long-range search overhead, and space amplification factor all quickly grow to the level of the STCS.
[0059] This embodiment uses the Bloom filter and fence pointer of the existing LSM tree to optimize the query, but improves the memory allocation mechanism of the Bloom filter. In order to reduce the overall false positive probability of the system under the same memory overhead, it is necessary to reasonably allocate the memory size of the Bloom filter for each layer. Since the fence pointer is adopted, the I / O overhead of probing and accessing the SSTable for each layer is the same, and the queried keys are very likely to be in the bottom layer. Therefore, less memory is allocated to the Bloom filter in the bottom layer and allocated to the upper layer to reduce the false positive probability of the upper layer, thereby reducing the overall query overhead of the system.
[0060] In one implementation, the method for determining the first target layer includes:
[0061] Step S101: For each non-lowest layer in the key-value storage system, when writing the last ordered string table, set this layer as the first target layer.
[0062] Specifically, the compression trigger condition in this embodiment is also different from that of the existing key-value storage system based on the LSM tree. In the existing key-value storage system based on the LSM tree, the merge and compression operation is generally triggered when the last SSTable is full. In this embodiment, for each layer except the lowest layer in the key-value storage system, the merge and compression operation is triggered when writing the last SSTable.
[0063] In one implementation, step S200 specifically includes:
[0064] Step S201: Define the first threshold as the product of the first quantity and the total file size of the first target layer;
[0065] Step S202: When there is an ordered string table in the second target layer whose file size is greater than or equal to the first threshold, determine whether there is an active ordered string table in the second target layer, and determine the active ordered string table from the ordered string tables whose file size is greater than or equal to the first threshold.
[0066] In one implementation, step S200 further includes:
[0067] Step S203: When there is no ordered string table in the second target layer whose file size is greater than or equal to the first threshold, determine that there is no active ordered string table in the second target layer.
[0068] Specifically, assume that the first target layer is layer Li and the second target layer is layer Li+1. Then the active SSTables in layer Li+1 are selected from the set of SSTables in this layer whose file size is not less than K * size(Li), where size(Li) is the total file size of layer Li, generally the last SSTable. If there is no SSTable in this layer whose file size is not less than K * size(Li), it is determined that there is no active SSTable in this layer.
[0069] In one implementation, step S300 specifically includes:
[0070] Step S301: Define the second threshold as the product of the second quantity and the total file size of the first target layer;
[0071] Step S302: When there is an ordered string table in the second target layer whose file size is greater than or equal to the second threshold, determine whether there is an active ordered string table in the second target layer, and determine the active ordered string table from the ordered string tables whose file size is greater than or equal to the second threshold.
[0072] In one implementation, step S300 further includes:
[0073] Step S303: When there is no ordered string table in the second target layer whose file size is greater than or equal to the second threshold, determine that there is no active ordered string table in the second target layer.
[0074] Specifically, if the second target layer is the bottom layer, the first target layer is layer Lk. The active SSTables in the bottom layer are selected from the set of SSTables in this layer whose file size is not less than Z * size(Lk), where size(Lk) is the total file size of layer Lk. If there is no SSTable in this layer whose file size is not less than Z * size(Lk), it is determined that there is no active SSTable in this layer.
[0075] In one implementation, the first quantity is the difference between the ratio of the total file sizes of the second target layer and the first target layer and 1; the second quantity is 1.
[0076] Specifically, in this embodiment, a reasonable state is set as K = T - 1 and Z = 1, where T = size(Li + 1) / size(Li). At this time, the target key query overhead, long-range search overhead, and space amplification factor are roughly the same as those of LCS, the short-range search overhead is also close to LCS, but the overall update overhead is as low as that of STCS.
[0077] Next, the impacts brought by adopting the method provided by the present invention in the state of K = T - 1 and Z = 1 will be elaborated in detail:
[0078] 1) Impact of LLCS on target key query overhead:
[0079] Target key queries are divided into zero-result queries and existence-result queries. A zero-result query means that the target key does not exist in the database and the entire database needs to be scanned layer by layer, with a relatively high overhead. However, after adopting the memory allocation design of the optimized Bloom filter, the overhead R of the zero-result query is described by the following formula (1), where M represents the memory size, N represents the number of key-value pairs, and T represents the number of SSTables in each layer.
[0080]
[0081] It can be seen that for any given value of T, R is a very small constant. Therefore, the complexity of the query overhead is O(e -M / N ), which is the same as the overhead of using LCS. However, LLCS reduces a large number of compression operations. The existence of a result query indicates that the target key exists in the database. In the worst case, the target key appears in the bottom layer. Therefore, the expected number V of file block reads is 1 for the bottom layer plus the sum of all Bloom filter false positives in the upper layers, as shown in Equation (2), where T is the number of SSTables in one layer in the upper layer, and p i represents the probability of a Bloom filter false positive in the i-th layer.
[0082]
[0083] Obviously, the complexity of its overhead is O(1) because is less than 1 at any time.
[0084] 2) Influence of LLCS on range queries:
[0085] Range queries need to scan all SSTables, which are divided into two categories: short-range queries and long-range queries. Short-range queries initiate at most O(T) file block read operations in the [0, L - 1] layers and one file block read operation in the bottom layer. Therefore, the complexity of the I / O overhead is O(1+(L - 1)*T). It should be noted that the I / O overhead increases with the increase of the T value. However, the value range of T is [1, T lim , where T lim is the upper limit of T. When T lim converges to 1, the complexity of the I / O overhead is O(1), and at this time, LLCS degenerates into the form of a single ordered array in LCS. The overhead of long-range queries mainly comes from the operation of sequentially accessing the bottom layer, and the complexity of its I / O overhead is where s represents the key value range of the query and B represents the largest key value range. Therefore, in long-range queries, LLCS maintains the same query overhead as LCS while reducing a large number of redundant compression operations in the upper layer.
[0086] 3) Influence of LLCS on updates:
[0087] In LLCS, one update operation causes O(1) merge compression operations in each layer of the [1, L - 1] layers and O(T) merge compression operations in the L-th layer, that is, the bottom layer. Therefore, the overall number of merge compressions per key value pair is O(L + T). Compared with LCS, the update overhead of LLCS is significantly reduced in the worst case.
[0088] 4) Influence of LLCS on space amplification:
[0089] In the worst case, all key-value pairs in [1, L - 1] are updates to the key-value pairs existing in layer L. Since the upper layer ([1, L - 1]) only accounts for Therefore, the space magnification factor at this time is This is the same as that of LCS, but it avoids the redundant merge and compression operations in the upper layer.
[0090] As Figure 5 shown, when implementing the LLCS strategy in this embodiment, the SSTables in the upper layer of the LSM tree are disordered, that is, there is a situation of cross redundancy. When their number reaches t, they are merged and compressed with the active SSTables in the lower layer and then flushed to the lower layer; if the number of SSTables in the lower layer reaches t, continue to recursively process downwards.
[0091] The advantages of the present invention are as follows:
[0092] The motivation for designing LLCS in the present invention is to provide greater flexibility in compression control and better adaptation to changing data loads. Compared with STCS, LLCS has lower target key query overhead, short-range search overhead, long-range search overhead, and space magnification factor, and has the same update overhead. Compared with LCS, LLCS reduces the overall update overhead, has the same target key query overhead, long-range search overhead, and space magnification factor, and similar short-range search overhead. The present invention optimizes the key-value storage system based on the LSM tree at the design level, enabling the current system to have better scalability, that is, the situation of performance degradation due to the increase in data storage capacity will not occur.
[0093] Based on the above embodiments, the present invention also provides a key-value storage system with compressible merging, as Figure 6 shown, the key-value storage system is constructed based on the LSM tree, and the key-value storage system includes:
[0094] Several layers for storing key-value pairs, wherein each non-bottom layer includes a first number of ordered string tables, and the bottom layer includes a second number of the ordered string tables;
[0095] A compression control module for determining a first target layer to be compressed in the key-value storage system and determining a second target layer below the first target layer, wherein the first target layer is not the bottom layer;
[0096] When the second target layer is not the bottom layer, it is determined whether there is an active ordered string table in the second target layer according to the first number and the total file size of the first target layer;
[0097] When the second target layer is the bottom layer, determine whether the active ordered string table exists in the second target layer according to the second quantity and the total file size of the first target layer;
[0098] When it exists, after orderly merging all the ordered string tables of the first target layer, orderly merge them with the active ordered string table to obtain a first merged ordered string table; move the first merged ordered string table into the second target layer;
[0099] When it does not exist, after orderly merging all the ordered string tables of the first target layer, obtain a second merged ordered string table; move the second merged ordered string table into the second target layer.
[0100] Based on the above embodiments, the present invention further provides a terminal, and its principle block diagram can be as Figure 7 shown. The terminal includes a processor, a memory, a network interface, and a display screen connected through a system bus. Among them, the processor of the terminal is used to provide computing and control capabilities. The memory of the terminal includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the terminal is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it realizes the compression and merging method of the key-value storage system. The display screen of the terminal can be a liquid crystal display screen or an electronic ink display screen.
[0101] Those skilled in the art can understand that Figure 7 the principle block diagram shown in
[0102] merely shows the block diagram of some structures related to the solution of the present invention, and does not constitute a limitation on the terminal to which the solution of the present invention is applied. A specific terminal may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0103] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided by the present invention can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0104] In summary, the present invention discloses a compression and merging method and system for key-value storage. The method is applied to a key-value storage system constructed by an LSM tree. The key-value storage system includes several layers for storing key-value pairs. Among them, each non-bottom layer includes a first number of ordered string tables, and the bottom layer includes a second number of the ordered string tables. The method includes: determining a first target layer to be compressed in the key-value storage system, and determining a second target layer located below the first target layer, where the first target layer is not the bottom layer; when the second target layer is not the bottom layer, judging whether there is an active ordered string table in the second target layer according to the first number and the total file size of the first target layer; when the second target layer is the bottom layer, judging whether there is the active ordered string table in the second target layer according to the second number and the total file size of the first target layer; when there is, after orderly merging all the ordered string tables of the first target layer, then orderly merging with the active ordered string table to obtain a first merged ordered string table; moving the first merged ordered string table into the second target layer; when there is no such table, after orderly merging all the ordered string tables of the first target layer, obtaining a second merged ordered string table; moving the second merged ordered string table into the second target layer. Since the overheads of target key query and long-range query mainly come from the bottom layer, and the space amplification factor mainly depends on the bottom layer, most of the compression operations on the upper layer have little effect on performance improvement. To significantly improve performance, it is necessary to change the compression strategy of the bottom layer. The present invention differentiates the compression mechanisms of the upper layer and the bottom layer, and introduces two parameters to respectively determine the active ordered string tables in the upper layer and the bottom layer. During merge compression, all the ordered string tables of the current layer and the active ordered string tables of the next layer are merged and compressed and then flushed to the next layer. By changing the specific values of the two parameters, the present invention can switch between different merge compression strategies, providing greater flexibility in compression control, expanding the design space of the LSM tree, and thus can better adapt to changing data loads. It solves the problem in the prior art that it is difficult to make a reasonable trade-off among write overhead, read overhead, and storage space overhead in the LSM tree compression strategy, resulting in the lack of scalability of the key-value storage system.
[0105] It should be understood that the application of the present invention is not limited to the above examples. For those of ordinary skill in the art, improvements or transformations can be made according to the above description. All such improvements and transformations should fall within the protection scope of the appended claims of the present invention.
Claims
1. A compression and merging method for a key-value storage system, characterized in that, The method is applied to a key-value storage system for constructing an LSM tree. The key-value storage system includes several layers for storing key-value pairs. Among them, each non-bottom layer includes a first number of ordered string tables, and the bottom layer includes a second number of the ordered string tables. The first number is the difference between the ratio of the total file sizes of the second target layer and the first target layer and 1, and the second number is 1. The method includes: Determine a first target layer to be compressed in the key-value storage system, and determine a second target layer located below the first target layer, where the first target layer is not the bottom layer; When the second target layer is not the bottom layer, judge whether there is an active ordered string table in the second target layer according to the first number and the total file size of the first target layer; When the second target layer is the bottom layer, judge whether there is the active ordered string table in the second target layer according to the second number and the total file size of the first target layer; When there is one, after orderly merging all the ordered string tables of the first target layer, and then orderly merging with the active ordered string table, obtain a first merged ordered string table; Move the first merged ordered string table to the second target layer; When there is none, after orderly merging all the ordered string tables of the first target layer, obtain a second merged ordered string table; Move the second merged ordered string table to the second target layer.
2. The compression and merging method of the key-value storage system according to claim 1, wherein The method for determining the first target layer includes: For each non-bottom layer in the key-value storage system, when writing the last ordered string table, use this layer as the first target layer.
3. The compression and merging method of the key-value storage system according to claim 1, characterized in that, Judging whether there is an active ordered string table in the second target layer according to the first number and the total file size of the first target layer includes: Define a first threshold as the product of the first number and the total file size of the first target layer; When there is an ordered string table in the second target layer whose file size is greater than or equal to the first threshold, judge that there is an active ordered string table in the second target layer, and determine the active ordered string table from the ordered string tables whose file size is greater than or equal to the first threshold.
4. The compression and merging method of the key-value storage system according to claim 3, wherein The method further includes: When there is no ordered string table in the second target layer whose file size is greater than or equal to the first threshold, judge that there is no active ordered string table in the second target layer.
5. The compression and merging method of the key-value storage system according to claim 1, characterized in that Judging whether there is the active ordered string table in the second target layer according to the second number and the total file size of the first target layer includes: Define a second threshold as the product of the second number and the total file size of the first target layer; When there is an ordered string table in the second target layer whose file size is greater than or equal to the second threshold, judge that there is an active ordered string table in the second target layer, and determine the active ordered string table from the ordered string tables whose file size is greater than or equal to the second threshold.
6. The compression and merging method of the key-value storage system according to claim 5, characterized in that, The method further includes: When there is no ordered string table with a file size greater than or equal to the second threshold in the second target layer, it is determined that there is no active ordered string table in the second target layer.
7. A key-value storage system that can be compressed and merged, characterized in that The key-value storage system is constructed based on the LSM tree. The key-value storage system includes: Several layers for storing key-value pairs. Among them, each non-bottom layer includes an ordered string table with a first quantity, and the bottom layer includes an ordered string table with a second quantity; A compression control module for determining a first target layer to be compressed in the key-value storage system and determining a second target layer below the first target layer. Among them, the first target layer is not the bottom layer, the first quantity is the difference between the ratio of the total file size of the second target layer and the first target layer and 1, and the second quantity is 1; When the second target layer is not the bottom layer, it is determined whether there is an active ordered string table in the second target layer according to the first quantity and the total file size of the first target layer; When the second target layer is the bottom layer, it is determined whether there is the active ordered string table in the second target layer according to the second quantity and the total file size of the first target layer; When there is, after the ordered string tables of the first target layer are merged in order, they are then merged in order with the active ordered string table to obtain a first merged ordered string table; the first merged ordered string table is moved into the second target layer; When there is no, after the ordered string tables of the first target layer are merged in order, a second merged ordered string table is obtained; the second merged ordered string table is moved into the second target layer.
8. A computer-readable storage medium having a plurality of instructions stored thereon, characterized in that, The instruction is suitable for being loaded and executed by a processor to implement the steps of the compression and merging method of the key-value storage system according to any one of claims 1-6 above.
Citation Information
Patent Citations
Memory system including key-value store
CN102929793A
Method and device for file compaction in KV (Key-Value)-Store system
CN106407224A