Cache-based Garbage Collection Optimization Method and Device for ZNS-SSD Key-Value Storage System

By caching SST files that are about to participate in compaction operations and adjusting the read and write process, the problem of low garbage collection efficiency in the ZNS-SSD key-value storage system is solved, and more efficient garbage collection and read and write performance is achieved.

CN120066984BActive Publication Date: 2025-07-08HUAQIAO UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510526278.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-07-08
Estimated Expiration
2045-04-25

AI Technical Summary

Technical Problem

During the garbage collection process of the existing ZNS-SSD key-value storage system, the migrated data quickly became invalid, resulting in an increase in read and write overhead, and the compaction operation is high, which affects system performance.

Method used

By selecting the Zone with invalid data proportion >70% as candidate zones, the valid SST files that are about to participate in the compaction operation are cached, and the cache status is recorded in the Manifest file, the read and write operation process is adjusted, and the GC strategy is adaptively adjusted to reduce the amount of data migration.

Benefits of technology

The cascading amplification brought by the combination of ZNS-SSD and LSM-Tree is reduced, the GC operation efficiency is improved, the write amplification of data migration is reduced, the read operation efficiency is improved, and the user-side write stagnation is avoided. It is suitable for placement strategies in different regions, and the GC performance is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120066984B_ABST
    Figure CN120066984B_ABST
Patent Text Reader

Abstract

The present invention discloses a garbage collection optimization method and device for a ZNS-SSD key-value storage system based on cache, which relates to the field of computer storage. The method includes the following steps: a selection step of selecting a Zone with an invalid data ratio > 70% as a candidate Zone; a judgment step of performing N rounds of judgment according to the judgment round value N; in each round of judgment, caching the valid SST files to be involved in the compaction operation in the candidate Zone into the memory; a garbage collection step of selecting the Zone with the largest ratio of invalid data plus cacheable SST data value in the candidate Zone for garbage collection operation; an optimization step of judging whether the amount of data cached in this GC operation is > 70% of the valid data. If so, directly end the garbage collection operation; otherwise, set the increased judgment round value to N + 1 and end the garbage collection operation. The present invention reduces the cascading amplification brought by the combination of ZNS-SSD and LSM-Tree, improves the efficiency of the GC operation, and reduces the write amplification of data migration.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer storage, and particularly to a garbage collection optimization method and device for a key-value storage system based on cache in ZNS-SSD. Background Art

[0002] ZNS-SSD is a new type of SSD technology that optimizes storage performance and extends device lifespan. By introducing the concept of partitioned storage and adopting a partitioned and sequential data writing and management method, it significantly improves storage efficiency, reduces write amplification, and effectively extends the device's service life.

[0003] The sequential write limitation of ZNS-SSD naturally fits with the append write and hierarchical merge methods adopted by the LSM-Tree (Log-Structured MergeTree) in the key-value storage system. The data structure of LSM-Tree exhibits excellent write performance and data persistence on ZNS-SSD. Combining ZNS-SSD with LSM-Tree can achieve significant performance improvement.

[0004] When the capacity of a certain Zone in ZNS-SSD is exhausted or reaches a preset threshold, a zone cleaning operation is triggered. During the cleaning process, the system will preferentially select a Zone with a lower proportion of invalid data for garbage collection (GC). During cleaning, the still valid data in the Zone will be migrated to a new target Zone, and the invalid data in the old Zone will be marked as recyclable. After the data migration is completed and there is no valid data in the old Zone, the system will perform a Reset operation to release the space of the Zone.

[0005] The migrated SST file may participate in the compaction operation of LSM-Tree subsequently. If the SST file triggers compaction immediately after the GC operation, it will cause the migrated data to become invalid quickly, resulting in read and write overhead. Some studies propose replacing the data migration in GC with a merge operation, but this method also brings new challenges: the complexity of the compaction operation is higher than simple data migration, which may increase the processing overhead. During the zone cleaning process, the compaction operation may cause the write to the user side to stall, affecting the overall performance of the system.

[0006] Therefore, it is necessary to design a garbage collection optimization method for a key-value storage system based on cache in ZNS-SSD to improve the efficiency of garbage collection operations and reduce the read and write amplification of this operation. Summary of the Invention

[0007] The purpose of the present invention is to solve the problems in the prior art.

[0008] The technical solution adopted by the present invention to solve its technical problems is to provide a garbage collection optimization method for a cache-based ZNS-SSD key-value storage system, including the following steps:

[0009] Selection step: Select a Zone with an invalid data ratio > 70% as a candidate Zone;

[0010] Judgment step: According to the judgment round value N, perform N rounds of judgment; in each round of judgment, cache the valid SST files to be involved in the compaction operation in the candidate Zone into the memory; the initial default value of N is 1;

[0011] Garbage collection step: Select the Zone with the largest ratio of invalid data plus cacheable SST data value in the candidate Zones for garbage collection operations;

[0012] Optimization step: Judge whether the amount of data cached in this GC operation is > 70% of the valid data. If so, directly end the garbage collection operation; otherwise, set the judgment round value to N + 1 and end the garbage collection operation.

[0013] Preferably, the judgment step includes the following steps:

[0014] Round judgment step: Confirm whether the current round of judgment is the first round of judgment. If so, enter the initial round judgment step; otherwise, enter the other round judgment step;

[0015] Initial round judgment step: Cache the valid SST files that satisfy Score x > 1 and Score x+1 to Score i all greater than 0.8 into the memory; where x is a certain level among levels 0 to i - 1, and i is the level where the valid SST file is located;

[0016] Other round judgment step: Cache the valid SST files with Score i > 0.8 into the memory;

[0017] Among them, , represents the size of the SST file data stored in the i-th layer, represents the size of the SST file data that can be stored in the i-th layer limit. The compaction operation threshold of the i-th layer is represented by Score i . Score i > 1 indicates that the i-th layer will actively trigger the compaction operation; Score i > 0.8 indicates that the i-th layer may trigger the compaction operation due to new data written in the upper layer.

[0018] Preferably, in each round of judgment, the valid SST files to participate in the compaction operation in the candidate Zone are cached in the memory, and it further includes: adding a cache field to the SST file metadata saved in the Manifest file. When caching an SST file, the cache field of the SST file metadata needs to be modified, writing a cache record in the Manifest file, marking the data in the Zone as cacheable SST, and adjusting the read operation process and the compaction operation process.

[0019] Preferably, for the adjustment of the read operation process and the compaction operation process, the adjustment to the read operation is: before performing a read operation on an SST file, first determine whether the SST file exists in the memory through the cache field. If it exists, perform the read operation from the memory.

[0020] Preferably, for the adjustment of the read operation process and the compaction operation process, the adjustment to the compaction operation is: when an SST file in the cache participates in the compaction operation, the corresponding cache record needs to be deleted, and the cached SST file data needs to be deleted.

[0021] The present invention also provides a garbage collection optimization device for a ZNS-SSD key-value storage system based on caching, which is used to implement the garbage collection optimization method for the ZNS-SSD key-value storage system based on caching described in any one of the above, including:

[0022] A selection module that selects a Zone with an invalid data ratio > 70% as a candidate Zone;

[0023] A judgment module that makes N rounds of judgments according to the judgment round value N; in each round of judgment, the valid SST files to participate in the compaction operation in the candidate Zone are cached in the memory; the initial default value of N is 1;

[0024] A garbage collection module that selects the Zone with the largest ratio of invalid data plus cacheable SST data value in the candidate Zone for garbage collection operation;

[0025] An optimization module that judges whether the amount of data cached in this GC operation is > 70% of the valid data. If so, directly end the garbage collection operation; otherwise, set the judgment round value to N + 1 and end the garbage collection operation.

[0026] The present invention has the following beneficial effects:

[0027] (1) The present invention reduces the cascading amplification brought by the combination of ZNS-SSD and LSM-Tree. At the same time, it improves the efficiency of the GC operation and reduces the write amplification of data migration;

[0028] (2) By adding new SST file metadata fields and writing cache records in the Manifest file, the present invention improves the read operation efficiency and solves the problem of recovery after a failure of the SST file in the cache;

[0029] (3) By adjusting the traditional GC strategy and selecting the Zone with the largest proportion of invalid data and cacheable SST, the present invention can effectively reduce the migration amount of valid data during the GC operation and speed up the GC operation;

[0030] (4) The present invention adjusts the judgment condition for caching SST in this solution according to the amount of cached data during the GC operation, making this solution more widely applicable and capable of effectively improving GC performance in different region placement strategies;

[0031] (5) The present invention performs a caching operation on the SST, avoiding the problem of write stagnation at the user end caused by frequently performing compaction operations on the SST files in the disk instead of the GC operation.

[0032] The following further elaborates on the present invention in detail with reference to the drawings and embodiments, but the present invention is not limited to the embodiments. Description of the Drawings

[0033] Figure 1 It is a method step diagram of the garbage collection optimization method for the cache-based ZNS-SSD key-value storage system according to the embodiment of the present invention;

[0034] Figure 2 It is a system structure diagram of the garbage collection optimization method for the cache-based ZNS-SSD key-value storage system according to the embodiment of the present invention;

[0035] Figure 3 It is a flow schematic diagram of the garbage collection optimization method for the cache-based ZNS-SSD key-value storage system according to the embodiment of the present invention;

[0036] Figure 4 It is a read operation flow schematic diagram of the garbage collection optimization method for the cache-based ZNS-SSD key-value storage system according to the embodiment of the present invention;

[0037] Figure 5 It is a compaction operation flow schematic diagram of the garbage collection optimization method for the cache-based ZNS-SSD key-value storage system according to the embodiment of the present invention;

[0038] Figure 6 It is a structure schematic diagram of the garbage collection optimization device for the cache-based ZNS-SSD key-value storage system according to the embodiment of the present invention. Detailed Embodiments

[0039] The present invention provides a garbage collection optimization method and device for a cache-based ZNS-SSD key-value storage system. A candidate Zone for GC operation is selected based on the proportion of invalid data. SST files that can be cached are marked according to the Score value of the layer and the CP pointer within the layer. SST files that are about to undergo subsequent compaction operations during the GC process are cached, valid data is migrated, and the Zone space is reclaimed. According to the amount of data cached in each GC operation, the judgment round of SST files that can be cached is adaptively adjusted. Except for the first round, the SST files that can be cached are directly judged according to the CP pointer of this layer. When the SST file is the SST file that the CP pointer can point to after increasing the round and the Score value of this layer is close to 1, this SST file is cached, and when performing the compaction operation, it is checked whether there is still the SST file cached in this layer in the cache. If so and the Score value of this layer is close to 1, the next compaction operation of this layer is triggered in advance.

[0040] See Figure 1 As shown, it is the method step diagram of the garbage collection optimization method for the cache-based ZNS-SSD key-value storage system in the embodiment of the present invention, including the following steps:

[0041] S101, Selection step, select the Zone with the proportion of invalid data > 70% as the candidate Zone;

[0042] S102, Judgment step, according to the judgment round value N, perform N rounds of judgment; in each round of judgment, cache the valid SST files to be involved in the compaction operation in the candidate Zone into the memory;

[0043] S103, Garbage collection step, select the Zone with the largest proportion of invalid data plus cacheable SST data value in the candidate Zone for garbage collection operation;

[0044] S104, Optimization step, judge whether the amount of data cached in this GC operation > 70% of the valid data. If so, directly end the garbage collection operation; otherwise, set the judgment round value to N + 1 and end the garbage collection operation.

[0045] Among them, the initial default value of N is 1.

[0046] Specifically, in the S102 judgment step, for the SST files that need data migration, if the file will participate in the next compaction operation, cache this SST file into the memory.

[0047] Specifically, the judgment between layers is through judgment, ; sum iRepresents the data size of the SST file stored in layer i, M i Represents the data size of the SST file that can be stored restricted by layer i. When Score i is greater than 1, select layer i to trigger the compaction operation. If there are multiple layers greater than 1 and equal, select the higher layer for the compaction operation. When the Score values of layers 0 to i are all < 1, it means that the compaction operation will not be triggered in this round, and subsequent judgments on the SST file S involved in the GC operation will not be performed. When the Score value of layer 0 > 1 and the Score values of layers 1 to i - 1 are all close to 1, or the Score of layer i - 1 and layer i > 1. It means that the layer i where the SST file S is located in this round may trigger the compaction operation for subsequent judgment. First, judge whether the SST file S has an overlapping key range with the SST file T that is about to trigger the compaction operation in layer i - 1. If so, it means that S is about to undergo the compaction operation and S is cached; if there is no overlapping key range, the Score value of layer i is close to or greater than 1, and CPi (compaction Point) points to S, indicating that S is about to undergo the compaction operation and S is cached.

[0048] See Figure 2 As shown, when the free space in the SSD is less than a certain amount, the GC operation is triggered. Zone1 is selected as the zone for area cleaning, and there is a valid SST file in it. After this file is judged to be cacheable, modify the metadata cache field of this SST in the Manifest and write the cache record, then mark the SST file in the zone and cache this SST file. Then perform the GC operation on Zone1 and reset Zone1 to reclaim space.

[0049] Specifically, in the S102 judgment step, a cache field is added to the SST metadata saved in Manifest, and a record of the cached SST is written in the Manifest file for subsequent fault recovery. And the read operation process is adjusted to improve the read performance. Before caching the SST, the cache field of the SST metadata in Manifest is changed to 1 (a value of 0 for this field indicates that the SST file is not in the cache), and then a record MemOnlySST: {file_number, key_range, level, creation_time, cache_address} is written, where file_number represents the SST file number, key_range represents the key range of the SST file, level represents the level where the SST is located, creation_time represents the creation time of the SST file, and cache_address represents the location in memory. This record is used to recover the data saved on the SSD in combination with the WAL file after a fault.

[0050] Specifically, when performing a read operation, query whether the cache field of the SST metadata in the Manifest file is 1. If it is 1, read the SST file from memory; if it is 0, read the SST file from the SSD.

[0051] Specifically, when the amount of cached data < 70% of the valid data volume in the Zone during the GC operation, increase the number of rounds of SST caching judgment. However, since the structure of the LSM tree changes after one round of compaction, the judgment of the SST for the next round of compaction operation is relatively complex. Therefore, except for the first round, if the Score value of this layer is close to 1 and the CP pointer will point to this SST file after the increased number of rounds, then directly cache this SST. After performing the compaction operation on layer i, check whether there is still a cached SST of layer i in the cache. If there is a cached SST and the Score value of layer i is close to 1, then trigger the compaction operation of this layer in advance.

[0052] See Figure 3 As shown in the flowchart of the garbage collection optimization method for the cache-based ZNS-SSD key-value storage system, specifically as follows:

[0053] 201. When the free space in the SSD is insufficient, trigger the GC operation and jump to 202.

[0054] 202. Take all Zones where the proportion of invalid data < 70% as candidate Zones for subsequent judgment, and jump to 203.

[0055] 203. For the SST that needs to migrate valid data in the Zone, perform subsequent judgment and jump to 204. If all the valid data in the Zone has been judged, jump to 213 for GC operation.

[0056] 204. Judge whether the current judgment of the cacheable SST is the first round. If it is, jump to 205 to judge whether the SST file undergoes the compaction operation in the first round. If not, jump to 210 for the judgment of the i-th round.

[0057] 205. Judge the SST files existing in the Zone. If Score0 > 1 and Score1 to Scorei > 0.8, it indicates that the SST file may participate in the compaction operation, and jump to 207. If the conditions are not met, perform the judgment for another situation and jump to 206.

[0058] 206. Judge whether the SST file is another situation that triggers the compaction operation. If Scorei or Scorei - 1 > 1, it indicates that the SST file may participate in the compaction operation, and jump to 207. Otherwise, it indicates that the SST file cannot participate in the subsequent compaction operation, and jump to 203 to judge other SST files in the Zone.

[0059] 207. If there is a key overlap range between the SST file T that triggers the compaction operation in the i - 1 layer and this SST file, it indicates that this SST is passively triggered for the compaction operation. Cache this SST file and jump to 209. Otherwise, perform the judgment for the condition of actively triggering the compaction operation and jump to 208.

[0060] 208. If this SST file is the file pointed to by the CPi pointer of layer i, it indicates that this file may actively trigger the compaction operation. Cache this SST file and jump to 209. Otherwise, it indicates that this SST file cannot participate in the subsequent compaction operation, and jump to 203 to judge other SST files in the Zone.

[0061] 209. Modify the metadata of this SST in the Manifest file and write the cache record. Cache this SST, mark the SST file in the SSD as a cacheable SST, and jump to 203 to judge other SST files.

[0062] 210. If after adaptive adjustment, the GC operation needs to judge the i-th round except the first round, jump to 211 to judge the SST file. If there is no additional judgment round, jump to 203 to judge other SST files.

[0063] 211. Whether the Score value of the layer where the valid SST file is located in the Zone is > 0.8. If it is greater, it indicates that a subsequent compaction operation may occur in this layer, and jump to 212. If it is < 0.8, it indicates that the probability of a subsequent compaction operation in this layer is small, and jump to 203.

[0064] 212. Whether the CP pointer of the layer where the valid SST file is located in the Zone can point to this SST after i compaction operations. If it can, it indicates that this SST file will trigger a compaction operation after i rounds of compaction. Cache this SST file and jump to 209. Otherwise, jump to 203.

[0065] 213. After all candidate Zones are judged, select the Zone with the most invalid data and cacheable data as the object of this GC operation. After caching the cacheable SST files and migrating the valid data, reset the Zone to reclaim space and jump to 214.

[0066] 214. Judge whether the amount of data cached in this GC operation is > 70% of the valid data. If it is greater, it indicates that this idea matches well with the placement method, and there is no need to increase the number of judgment rounds for cacheable SSTs, and jump to 216. Otherwise, jump to 215.

[0067] 215. The amount of data in the cached SST file < 70% of the valid data, indicating that it matches poorly with the current placement method. Increase the number of judgment rounds for cacheable SSTs, increase the amount of cacheable data to improve the performance of the present invention, and jump to 216.

[0068] 216. Complete the GC operation.

[0069] See Figure 4 As shown in the following figure, it is the adjusted read operation flow chart, specifically:

[0070] 301. Perform a read operation on the SST file.

[0071] 302. Search the Memtable to determine whether the SST file is in the Memtable. If it is, jump to 306. If not, jump to 303.

[0072] 303. Search the Immutable Memtable to determine whether the SST file is in the Immutable Memtable. If it is, jump to 306. If not, jump to 304.

[0073] 304. Search for the location of this SST in the LSM tree and jump to 305.

[0074] 305. Search for the metadata of the SST file in the Manifest file. If the cache field is 1, it means the SST is cached and jump to 306; if it is 0, it means the file is not cached and jump to 307.

[0075] 306. Read the SST file from memory and jump to 308.

[0076] 307. If the SST file is not in memory, read the corresponding block of the SST file from the SSD and jump to 308.

[0077] 308. The read operation ends.

[0078] See Figure 5 As shown in the adjusted compaction operation flow chart below:

[0079] 401. The LSM tree triggers the compaction operation and jumps to 402.

[0080] 402. Determine whether the SST file is cached through the Manifest file. If it is cached, jump to 403; otherwise, jump to 404.

[0081] 403. Read the SST file participating in the compaction operation from memory and jump to 405.

[0082] 404. Read the corresponding block of the SST file from the SSD for the subsequent compaction operation and jump to 406.

[0083] 405. After the compaction operation, write back the new data, delete the corresponding cache record in the Manifest file, and the SST file data saved in memory, and jump to 407.

[0084] 406. After the compaction operation, write back the new data, mark the corresponding SST file on the SSD as invalid data to facilitate subsequent GC operations, and jump to 407.

[0085] 407. Determine whether there is a cached SST file in the same level as the currently triggered compaction operation in memory. If there is, jump to 408; if not, jump to 510.

[0086] 408. Determine whether the Score value of this layer is greater than 0.8. If it is greater than 0.8, jump to 409; if it is less than 0.8, do not trigger the compaction operation in advance and jump to 510.

[0087] 409. Trigger the compaction operation actively triggered by the cached SST file in advance.

[0088] 410. End this compaction operation.

[0089] Specifically, the method of the embodiment of the present invention is experimentally verified. Use db_bench to randomly write 26,488,000 key-value pairs with a key size of 16B and a value size of 1000B to the ZNS-SSD device, and then use db_bench to perform overwrite update operations on the key-value pairs to obtain relevant data: The number of SST files migrated by the GC operation is 289, of which 63 SST files are cacheable SST files in the initial round of judgment, and there are a few SST files among the 63 SST files that trigger the GC migration operation multiple times during the compaction process for data migration. There are 4 SST files that are cacheable SST files in other rounds. By caching the cacheable SST files of the present invention, the write amplification caused by the GC migration of valid data and the compaction operation can be reduced, and the data migration operation can be reduced by about 22%. After optimizing the read process and the compaction process, the number of read and write I / Os of the system can be reduced, and the read and write latency can be reduced.

[0090] See Figure 6 As shown, it is a schematic structural diagram of a garbage collection optimization device for a cache-based ZNS-SSD key-value storage system according to an embodiment of the present invention, including:

[0091] A selection module 601 that selects a Zone with an invalid data ratio > 70% as a candidate Zone;

[0092] A judgment module 602 that makes N rounds of judgments according to the judgment round value N; In each round of judgment, the valid SST files to participate in the compaction operation in the candidate Zone are cached in the memory; The initial default value of N is 1;

[0093] A garbage collection module 603 that selects the Zone with the largest ratio of invalid data plus cacheable SST data value in the candidate Zone for garbage collection operations;

[0094] An optimization module 604 that determines whether the amount of data cached in this GC operation is > 70% of the valid data. If so, directly end the garbage collection operation; Otherwise, set the judgment round value to N + 1 and end the garbage collection operation.

[0095] The module function implementation of a garbage collection optimization device for a cache-based ZNS-SSD key-value storage system according to an embodiment of the present invention is the same as the method of a garbage collection optimization device for a cache-based ZNS-SSD key-value storage system according to an embodiment of the present invention, which will not be elaborated here.

[0096] It can be seen that the present invention provides a garbage collection optimization method and device for a cache-based ZNS-SSD key-value storage system. By caching the SST files that are about to undergo compaction operations, it avoids the cascading amplification caused by triggering compaction operations after GC operations migrate SST files. By adjusting the GC strategy, the cache performance is optimized, and an adaptive cacheable SST determination condition is adopted so that the present invention can play a role under different placement strategies.

[0097] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A garbage collection optimization method for a cache-based key-value storage system of ZNS-SSDs, characterized in that, It includes the following steps: Selection step: Select the Zone with the invalid data ratio > 70% as the candidate Zone; Judgment step: According to the judgment round value N, perform N rounds of judgment; In each round of judgment, cache the valid SST files to be involved in the compaction operation in the candidate Zone into the memory; the initial default value of N is 1; Garbage collection step: Select the Zone with the largest ratio of invalid data plus cacheable SST data value in the candidate Zones for garbage collection operation; Optimization step: Judge whether the data volume cached in this GC operation is > 70% of the valid data. If so, directly end the garbage collection operation; Otherwise, set the judgment round value to N + 1 and end the garbage collection operation; The said judgment step includes the following steps: Round judgment step: Confirm whether the current round of judgment is the first round of judgment. If so, enter the initial round judgment step; otherwise, enter the other round judgment steps; Initial judgment step, caching valid SST files that satisfy Score x > 1 and Score x+1 to Score i all greater than 0.8 into memory; where x is a certain level among the 0 to i - 1 levels, and i is the level where the valid SST file is located; Other round judgment steps, cache the valid SST files with Score i > 0.8 into memory; Among them, sum i represents the data size of the SST file stored in the i-th layer, M i represents the data size of the SST file that can be stored restricted by the i-th layer, through Score i indicates the compaction operation threshold triggered by the i-th layer, Score i > 1 indicates that the i-th layer will actively trigger the compaction operation; Add a cache field to the SST file metadata saved in the Manifest file. When caching an SST file, modify the cache field of the SST file metadata, write a cache record in the Manifest file, mark the data in the Zone as cacheable SST, and adjust the read operation process and the compaction operation process.

2. The garbage collection optimization method for the cache-based ZNS-SSD key-value storage system according to claim 1, wherein The said adjustment of the read operation process and the compaction operation process: The adjustment of the read operation is: Before reading an SST file, determine whether the SST file exists in the memory through the cache field. If it exists, perform the read operation from the memory.

3. The garbage collection optimization method for the cache-based ZNS-SSD key-value storage system according to claim 1, wherein The said adjustment of the read operation process and the compaction operation process: The adjustment of the compaction operation is: After a cached SST file participates in the compaction operation, delete the corresponding cache record and delete the cached SST file data.

4. A garbage collection optimization device for a cache-based ZNS-SSD key-value storage system, characterized in that, Used to implement the garbage collection optimization method of the cache-based ZNS-SSD key-value storage system described in any one of claims 1 to 3, including: Selection module: Select the Zone with the invalid data ratio > 70% as the candidate Zone; Judgment module: According to the judgment round value N, perform N rounds of judgment; in each round of judgment, cache the valid SST files to be involved in the compaction operation in the candidate Zone into the memory; the initial default value of N is 1; Garbage collection module: Select the Zone with the largest ratio of invalid data plus cacheable SST data value in the candidate Zones for garbage collection operation; Optimization module: Judge whether the data volume cached in this GC operation is > 70% of the valid data. If so, directly end the garbage collection operation; otherwise, set the judgment round value to N + 1 and end the garbage collection operation.

Citation Information

Patent Citations

  • SSD-based key-value separation storage method supporting efficient storage space management

    CN112131140A

  • Method and system for reducing garbage collection and write amplification of key-value separation storage system

    CN112395212A