Garbage recycling method and device and computing device cluster

By garbage collection based on the amount of garbage data in the blob file, the problem of inaccurate garbage cleaning in the existing technology is solved, and a more efficient and accurate garbage collection effect is achieved.

CN120162276APending Publication Date: 2025-06-17HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410294676.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-12-14
Filing Date
2024-03-14
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

The prior art is based on the writing time in garbage collection, which can easily lead to inaccurate garbage cleaning, resulting in problems such as amplification of writes or untimely cleaning of garbage data.

Method used

By counting the amount of junk data in the blob file, we can determine whether the recycling conditions are met (the amount of junk data is greater than or equal to the target value or ranked in the top N by size), and garbage collection is carried out based on this.

Benefits of technology

Improve the accuracy of garbage collection, avoid write amplification problems caused by invalid cleaning, and clean up hot garbage data in a timely manner.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120162276A_ABST
    Figure CN120162276A_ABST
Patent Text Reader

Abstract

The invention provides a garbage collection method and device and a computing equipment cluster, belongs to the technical field of databases, and aims to perform garbage collection on the basis of the garbage data volume of block files and collect the block files with the garbage data volume meeting the condition, so that the accuracy of database garbage collection can be improved. The method comprises the following steps: a key value storage system counts the junk data volume of a plurality of block files, wherein the junk data volume of the block files is used for reflecting the size of deleted and / or updated values in the block files; the key value storage system judges whether the junk data volume of each block file in the plurality of block files meets a recovery condition or not; the recovery condition is as follows: the junk data volume is greater than or equal to a target value, or the junk data volumes of the plurality of block files are ranked from large to small and then the first N block files are ranked; wherein N is an integer greater than 0; and the key value storage system recycles the target block file meeting the recycling condition.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims the priority of a Chinese patent application with the application number 202311722980.1 and the application title "A Garbage Collection Method, Device and Computing Equipment Cluster" submitted to the National Intellectual Property Administration on December 14, 2023, the entire content of which is incorporated herein by reference. Technical Field

[0002] This application relates to the field of database technologies, and in particular, to a garbage collection method, device and computing equipment cluster. Background Art

[0003] RocksDB is a storage engine suitable for databases, which adopts the log-structured merge tree (LSM Tree) architecture to store key-value (KV) data. BlobDB is a solution for RocksDB to handle large values. By storing large values in a dedicated blob file and only storing pointers corresponding to the values in the LSM-tree, it is possible to avoid copying the values repeatedly during the compaction process and reduce the resource consumption caused by write amplification.

[0004] For garbage requests of blob files, it is generally considered that the earlier written data is more likely to be garbage data. Therefore, related technologies usually perform garbage cleaning on blob files with a long write time, which is likely to cause inaccurate garbage cleaning. For example, the data in blob files with a long write time is not necessarily garbage data, and rewriting this data during garbage cleaning will cause additional write amplification. Another example is that only blob files with a long write time are garbage cleaned, while relatively new blob files are not garbage cleaned, which is likely to cause untimely garbage data cleaning. Summary of the Invention

[0005] This application provides a garbage collection method, device and computing equipment cluster. Garbage cleaning is performed based on the amount of garbage data in blob files, and blob files whose garbage data amount meets the conditions are recycled, which can improve the accuracy of database garbage collection.

[0006] To achieve the above object, the embodiments of this application provide the following technical solutions:

[0007] In a first aspect, a garbage collection method is provided, which is applied to a key-value storage system. In the key-value storage system, the keys and values of data are stored in different files respectively, and among them, the values of data are stored in the form of block files in the key-value storage system. The method includes: the key-value storage system counts the garbage data volume of multiple block files, and the garbage data volume of the block file is used to reflect the size of the deleted and / or updated values in the block file; the key-value storage system determines whether the garbage data volume of each block file among the multiple block files meets the recycling condition; the recycling condition is that the garbage data volume is greater than or equal to the target value, or the block files ranked in the top N in the order of the garbage data volume of the multiple block files from large to small; where N is an integer greater than 0; the key-value storage system recycles the target block files that meet the recycling condition.

[0008] As can be seen from the above, in the above garbage collection method, the key-value storage system performs garbage cleaning based on the garbage data volume of the blob files, and recycles the blob files whose garbage data volume meets the conditions. Compared with performing garbage cleaning based on the write time of the blob files, the garbage data volume can more intuitively and accurately reflect the specific garbage situation of the blob files, and is more in line with the actual situation of the blob files. Therefore, recycling the blob files whose garbage data volume meets the conditions can recycle the blob files that really need garbage cleaning, and at the same time avoid the write amplification problem caused by invalid cleaning, thereby improving the accuracy of garbage collection in the key-value storage system.

[0009] In a possible implementation manner, the key-value storage system determines whether the garbage data volume of each block file among the multiple block files meets the recycling condition, including: the key-value storage system sorts the multiple block files in the order of the garbage data volume from large to small to obtain a sorting result; if the block file is any one of the top N block files in the sorting result, the key-value storage system determines that the garbage data volume of the block file meets the preset condition; if the block file is not any one of the top N block files in the sorting result, the key-value storage system determines that the garbage data volume of the block file does not meet the preset condition.

[0010] As can be seen from the above, sorting the block files according to the size of the garbage data volume, the block file with a more forward position indicates that its garbage data volume is relatively larger. Therefore, screening out the block files with a forward position and performing garbage collection can improve the accuracy of garbage cleaning.

[0011] In a possible implementation, the key-value storage system determines whether the amount of garbage data in each of multiple block files meets the recycling condition, including: the key-value storage system respectively compares the amount of garbage data in each of the multiple block files with a target value. If the amount of garbage data in a block file is greater than or equal to the target value, it is determined that the amount of garbage data in the block file meets the recycling condition. If the amount of garbage data in a block file is less than the target value, it is determined that the amount of garbage data in the block file does not meet the recycling condition.

[0012] As can be seen from the above, by setting the target value of the amount of garbage data, blob files that meet the conditions can be screened out. Furthermore, block files with the amount of garbage data greater than or equal to the target value can be deleted in a timely manner, improving the timeliness of garbage cleaning.

[0013] In a possible implementation, the multiple block files are all the block files in the key-value storage system. After the key-value storage system recycles the target block file that meets the recycling condition, the method further includes: determining the hot block files in the key-value storage system with an access frequency greater than or equal to a preset frequency; recycling the hot block files that meet the recycling condition.

[0014] As can be seen from the above, after recycling the target blob file, the key-value storage system performs garbage cleaning again on the hot blob files. By specifically cleaning the hot garbage data, the purpose of timely cleaning the hot garbage can be achieved.

[0015] In a possible implementation, the multiple block files are the hot block files in the key-value storage system with an access frequency greater than or equal to a preset frequency.

[0016] As can be seen from the above, the key-value storage system only needs to execute the GC policy on the hot blob files in the key-value storage system, rather than executing the GC policy on all the blob files in the key-value storage system. The number of blob files targeted is reduced, thereby reducing the resource consumption during garbage collection and alleviating the pressure on the key-value storage system.

[0017] In a possible implementation, the key-value storage system recycles the target block file that meets the recycling condition, including: the key-value storage system writes the non-garbage values in the target block file to a block file outside the target block file in the key-value storage system, and deletes the target block file; the non-garbage values are the values in the target block file that have not been updated and / or deleted.

[0018] As can be seen from the above, when recycling the target blob file, the present application first copies the non-garbage values (value) in the second blob file to a new blob file, and then deletes the target blob file, which can ensure that only the garbage values (value) are deleted during garbage cleaning of the target blob file, while retaining the non-garbage values (value).

[0019] In a possible implementation, the method further includes: the key-value storage system updates the amount of garbage data in the block files in the key-value storage system; the key-value storage system counts the amount of garbage data in multiple block files, including: the key-value storage system counts the latest amount of garbage data in multiple block files.

[0020] As can be seen from the above, when the key-value storage system performs a background merge operation, the garbage data statistic of the blob file is continuously updated, that is, the key-value storage system updates the amount of garbage data in the blob file in the key-value storage system. To further improve the accuracy of garbage collection, the key-value storage system counts the latest amount of garbage data in multiple blob files.

[0021] In a possible implementation, the persistent storage table file in the key-value storage system is used to store the keys of the data and the address information of the values corresponding to the keys, and the address information is used to reflect the block file where the value is located; the key-value storage system updates the amount of garbage data in the block files in the key-value storage system, including: if the first key is updated or deleted, the key-value storage system updates the amount of garbage data in the first block file to the sum of the size of the value corresponding to the first key and the original amount of garbage data in the first block file; wherein, the first block file is the block file indicated by the address information of the value corresponding to the first key.

[0022] As can be seen from the above, once a certain key is updated or deleted, the key-value storage system of the present application will update the amount of garbage data in the blob file where the value corresponding to the key is located, making the amount of garbage data in the blob file more accurate, laying a foundation for subsequent garbage collection.

[0023] In a possible implementation, the key-value storage system updates the amount of garbage data in the first block file to the sum of the size of the value corresponding to the first key and the original amount of garbage data in the first block file, including: storing the identifier of the first block file and the size of the value corresponding to the first key into the meta-information of the persistent storage table file where the first key is located; after the key-value storage system completes the merge of all persistent storage table files, reading the meta-information of the persistent storage table file where the first key is located, and updating the amount of garbage data in the first block file to the sum of the size of the value corresponding to the first key and the original amount of garbage data in the first block file.

[0024] As can be seen from the above, updating the amount of garbage data in the blob file together after the merge is completed can reduce resource consumption.

[0025] In a possible implementation, the key-value storage system updates the amount of garbage data in the first block file to the sum of the size of the value corresponding to the first key and the original amount of garbage data in the first block file, including: during the process of merging the persistent storage table file where the first key is located in the key-value storage system, updating the amount of garbage data in the first block file to the sum of the size of the value corresponding to the first key and the original amount of garbage data in the first block file.

[0026] As can be seen from the above, when the key-value storage system performs file merging, when encountering an updated or deleted key, it will update the amount of garbage data in the corresponding blob file, so that the amount of garbage data in the blob file can be updated in a timely manner.

[0027] In a second aspect, a garbage collection device is provided, which is applied to a key-value storage system. In the key-value storage system, the keys and values of the data are stored in different files respectively, and the values of the data are stored in the form of block files in the key-value storage system. The device includes: an acquisition module for counting the amount of garbage data in multiple block files, and the amount of garbage data in the block file is used to reflect the size of the deleted and / or updated values in the block file; a judgment module for judging whether the amount of garbage data in each of the multiple block files meets the recycling condition; the recycling condition is: the amount of garbage data is greater than or equal to the target value, or the block files ranked in the first N block files after sorting the amount of garbage data in the multiple block files from largest to smallest; where N is an integer greater than 0;

[0028] A processing module for recycling the target block file that meets the recycling condition.

[0029] In a third aspect, a garbage collection system is provided. The system includes: a data storage device for providing data storage services. In the data storage device, the keys (keys) and values (values) of the data are stored in different files, and the blob files in the data storage device are used to store the values (values) of the data; a garbage collection device for: counting the amount of garbage data in multiple block files, and the amount of garbage data in the block file is used to reflect the size of the deleted and / or updated values in the block file; judging whether the amount of garbage data in each of the multiple block files meets the recycling condition; the recycling condition is: the amount of garbage data is greater than or equal to the target value, or the block files ranked in the first N block files after sorting the amount of garbage data in the multiple block files from largest to smallest; where N is an integer greater than 0; recycling the target block file that meets the recycling condition.

[0030] In a fourth aspect, a computing device cluster is provided, including at least one computing device, and each computing device includes a processor and a memory; the processor of at least one computing device is used to execute the instructions stored in the memory of at least one computing device, so that the computing device cluster executes the garbage collection method as described in the first aspect above.

[0031] In a fifth aspect, there is provided a computer program product including instructions which, when run on a cluster of computing devices, cause the cluster of computing devices to execute the garbage collection method as described in the first aspect above.

[0032] In a sixth aspect, there is provided a computer-readable storage medium including computer program instructions which, when executed by a cluster of computing devices, cause the cluster of computing devices to execute the garbage collection method as described in the first aspect above.

[0033] It should be noted that, for any possible implementation manner in each of the above aspects, combinations can be made on the premise that the solutions do not conflict with each other. Description of the Drawings

[0034] Figure 1 Schematic diagram of the writing principle of RocksDB provided by an embodiment of the present application;

[0035] Figure 2 Schematic diagram of the data writing process of RocksDB using BlobDB provided by an embodiment of the present application;

[0036] Figure 3 Schematic diagram of the GC module of BlobDB reclaiming blob files provided by an embodiment of the present application;

[0037] Figure 4 Schematic diagram of the process of a garbage collection method provided by an embodiment of the present application Figure 1 ;

[0038] Figure 5 Schematic diagram of the garbage collection principle provided by an embodiment of the present application;

[0039] Figure 6 Schematic diagram of the process of a garbage collection method provided by an embodiment of the present application Figure 2 ;

[0040] Figure 7 Schematic diagram of the structure of a garbage collection device provided by an embodiment of the present application;

[0041] Figure 8 Schematic diagram of the structure of a garbage collection system provided by an embodiment of the present application;

[0042] Figure 9 Schematic diagram of the structure of a computing device provided by an embodiment of the present application;

[0043] Figure 10 Schematic diagram of the structure of a computer cluster provided by an embodiment of the present application;

[0044] Figure 11Another structural schematic diagram of a computer cluster provided by an embodiment of the present application. Detailed implementation manners

[0045] The technical solutions provided by the present application will be described in detail below in conjunction with the accompanying drawings. Although some embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present application. It should be understood that the drawings and embodiments of the present application are only for exemplary purposes and are not used to limit the protection scope of the present application.

[0046] In the description of the embodiments of the present application, the term "including" and its similar terms should be understood as an open inclusion, that is, "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". Terms such as "first", "second", etc. may refer to different or the same objects. There may also be other explicit and implicit definitions hereinafter.

[0047] In the present application, "at least one" means one or more, and "a plurality" means two or more. "And / or" describes the association relationship of associated objects and indicates that three relationships may exist. For example, A and / or B may mean: including the case where A exists alone, A and B exist simultaneously, and B exists alone, where A and B may be singular or plural. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. "At least one (item)" or its similar expression refers to any combination of these items, including any combination of single item (item) or plural items (items). For example, at least one (item) of a, b, or c may mean: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, c may be single or multiple.

[0048] To make the technical solutions provided by the present application clearer, before specifically describing the technical solutions provided by the present application, some related terms and related technologies involved in the embodiments of the present application are first introduced.

[0049] Key-Value (KV) storage system: It is a non-relational database (NoSQL) that uses simple key-value pairs to store data. In a KV storage system, each data item consists of a unique key and a value associated therewith. These key-value pairs are usually stored in memory to provide very fast data access speed.

[0050] Log Structured Merge Tree (LSM Tree): A data structure used in key-value storage systems, mainly for optimizing disk I / O operations and improving write performance. The basic idea of the LSM Tree is to write write operations (including updates and deletes) first into a log file instead of directly into the main storage structure. In this way, all write operations can be performed sequentially and continuously, thus greatly improving write performance. When the log files grow to a certain extent, they are merged into the main storage structure, and this process is called the merge operation.

[0051] RocksDB: A KV storage system based on the Log Structured Merge Tree (LSM Tree). Among them, the architecture of RocksDB includes: MemTable, which is used to store recently written data; SSTable (Sorted String Table), which is used to store data that has been written to disk; Log File, also known as Write-Ahead Log (WAL), which is used to record all write operations; Compaction, which is used to merge multiple SSTables into a larger SSTable. When performing a write operation in RocksDB, the data is first written into the WAL and at the same time into the MemTable in memory. When the MemTable is full, it is converted into an ImmutableMemTable, and the data in the Immutable MemTable will eventually be written to the data file SSTable on persistent storage (such as disk). The data files on disk are divided into many layers. The higher the layer, the fewer the files, and the lower the layer, the more the files. When a certain layer is full, it will trigger the background Compaction thread to merge the data files in the upper layer to the lower layer. After the data is merged to the lower layer, the SSTable files in this layer can be deleted.

[0052] Merge (also known as combine): It can also be called compression, which is used to organize and merge existing records (data), thereby deleting some records that are no longer valid (for example, removing duplicate update or delete operation records), reducing the data scale and the number of files, and accelerating the reading speed. In the embodiments of this application, the merge process may include the Flush process (also known as minor compaction) and the Compaction process (also known as major compaction). Among them, the Flush process is to dump the Immutable MemTable in memory into the SST files of the 0th level (level, which can be simply referred to as layer or level) on the hard disk (that is, SSTable files). The Compaction process is to dump the SST files from a lower level to a higher level. The merge process in the following text mainly refers to the Compaction process. In the embodiments of this application, the 0th level on the hard disk refers to the organizational level of the SST files in the storage resource pool composed of the hard disk.

[0053] Junk data: Deleted data or old data before update.

[0054] Garbage Collection (GC): Garbage collection in a database mainly involves managing and cleaning up data that is no longer needed, obsolete, or invalid to ensure the efficient operation of the database and the effective utilization of storage space.

[0055] For ease of understanding, first, in combination with Figure 1 Introduce the write principle of RocksDB. Since RocksDB uses the LSM-Tree storage structure, to improve the data write performance, all data addition, deletion, and modification operations are written into the LSM-Tree in an append manner. The general process of writing data in RocksDB includes: data is first written into the MemTable, and at the same time, the write-ahead log records the data. When the MemTable is full, it becomes an Immutable MemTable, and then the Immutable MemTable will be flushed into the SST files of the L0 layer. At the same time, to balance the read performance, the data files in the Ln layer will be merged into the Ln+1 layer through the merge (Compaction) operation.

[0056] In the large KV scenario (that is, the size of the key (Key) or value (Value) exceeds the normal range), due to the large amount of data to be merged, Compaction will introduce a high write amplification and consume a large amount of resources such as the Central Processing Unit (CPU), memory, and Input / Output (IO). Therefore, BlobDB is introduced as a solution for processing large values.

[0057] BlobDB adopts a strategy of separate processing for keys and values. By storing large values in a dedicated block (blob) file and only storing their pointers in the LSM-tree, BlobDB aims to reduce the repeated copying of values during the compaction process, thereby reducing the write amplification.

[0058] In the embodiments of this application, the implementation of BlobDB can be an old implementation encapsulated on top of RocksDB or a new implementation directly embedded in RocksDB.

[0059] As Figure 2 shown, the data writing process of RocksDB that adopts BlobDB generally includes: adopting the method of separate storage for KV, putting the Value into the blob file separately, only merging and compressing the Key during compaction, and the Value is recycled through the GC module. Generally, in the large KV scenario, the Key is relatively small and the Value is large. Only the smaller Key and the address information of the corresponding Value are stored in the LSM-Tree. Therefore, the amount of data involved during the merge compaction (Compaction) is small, and thus the write amplification during Compaction can be reduced.

[0060] As Figure 3 shown, currently, the process of the GC module of BlobDB recycling the blob file can include the following stages:

[0061] (1) Metadata maintenance stage

[0062] cutoff: Represents the range of blob files that need to be recycled and can be updated before each Compaction.

[0063] oldest_blob_file_num: Records the number of the oldest blob file recorded in the SSTable file (sst->oldest_blob_file_num).

[0064] (2) Garbage marking stage

[0065] Compaction continuously iterates the Value corresponding to the Key. If the blob file number where the Value is located is within the cutoff, the Value will be rewritten into a new blob file. When the SSTable file is merged, the reference count of the corresponding blob file will be decremented by one.

[0066] (3) Garbage cleaning stage

[0067] When the reference count of a blob file before the cutoff drops to 0, it will be cleaned up by the GC module.

[0068] However, there are the following problems in the process of the GC module recycling blob files in the related art:

[0069] On the one hand, even if there is no garbage, the blob files before the cutoff will be rewritten repeatedly. That is, the existing GC strategy assumes that the earlier written data is garbage, so the older data will be rewritten repeatedly. However, in the actual business load, this is not the case. The older data is not necessarily garbage data, so this will cause additional write amplification.

[0070] On the other hand, the garbage after the cutoff cannot be cleaned up in time. Only when the blob files before the cutoff are cleaned up will it be the turn of the subsequent blob files. Therefore, even if the business generates garbage data later (such as deleting or updating the previous old data), it will not be cleaned up in time.

[0071] In view of the above problems, the embodiments of the present application provide a garbage collection method, device and computing device cluster, which perform garbage cleaning based on the amount of garbage data in the blob file, and recycle the blob files whose amount of garbage data meets the conditions, which can improve the accuracy of database garbage collection. For example, when RocksDB performs a background Compaction operation (i.e., merging data files), the garbage data statistics of the blob file will be continuously updated. The background GC module will sort all blob files according to the garbage data statistics at regular intervals, and according to the cleaning range set by the user or the default, clean up the blob files with the top-ranked garbage data amount. When there is no garbage data in the blob file, the GC module stops working.

[0072] The garbage collection method provided by the embodiments of the present application will be described in detail below with reference to the accompanying drawings. It should be noted that the embodiments of the present application can be borrowed or referenced from each other. For example, for the same or similar steps, the method embodiments, system embodiments and device embodiments can all be referenced from each other without limitation.

[0073] First, an example application scenario of the garbage collection method provided by the embodiments of the present application is introduced. The embodiments of the present application provide a garbage collection method that can be applied to a KV storage system. For example, it can be applied to a scenario where there are a large number of data deletions in a KV storage system based on the LSM-Tree structure. Among them, the KV storage system based on the LSM-Tree structure is, for example but not limited to, RocksDB, LevelDB, or BlobDB, etc. In this key-value storage system, the keys (Keys) and values (Values) of the data are stored in different files. For example, the Value can be separately placed in a blob file, and only the smaller Key and the address information of the corresponding Value are stored in the LSM-Tree. For example, the Key and the address information of the corresponding Value can be stored in an SST file.

[0074] In some embodiments, as shown in Figure 2 the external client can send a write request to the KV storage system to request access to the KV storage system (such as writing or deleting data, etc.). Further, in response to the user's write request, the storage engine of the KV storage system can drive the data to be stored in the storage layer to implement the user's write request. At the same time, the KV storage system can use a WAL to record the specific write operations. Specifically, the written data will first be stored in the MemTable in the memory. During the Flush process, the Key corresponding to the written data and the address information of the corresponding Value will be dumped into the SST file in the storage layer, and the Value will be dumped into the blob file in the storage layer. During the Compaction process, the SST files are dumped from a lower level to a higher level, and the blob files perform garbage collection. The specific garbage collection process can be referred to the introduction in the following method embodiments and will not be elaborated here. Among them, the storage layer can be any storage medium, such as a disk (Disk), an optical disc, etc.

[0075] It should be noted that the embodiments of the present application do not limit the specific application scenario. The system architecture and business scenarios described in the embodiments of the present application are to more clearly illustrate the technical solutions of the embodiments of the present application, and do not constitute a limitation to the technical solutions provided by the embodiments of the present application. Those of ordinary skill in the art know that with the evolution of the network architecture and the emergence of new business scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.

[0076] The following introduces the garbage collection method provided by the embodiments of the present application.

[0077] Figure 4 is a schematic flowchart of a garbage collection method provided by an embodiment of the present application. As shown in Figure 4 the method may include the following steps:

[0078] S101. The key-value storage system counts the amount of garbage data in multiple block files.

[0079] Among them, the amount of garbage data in the block file is used to reflect the size of the deleted and / or updated values in the block file. The multiple block files are some or all of the blob files in the key-value storage system.

[0080] In a possible design, the key-value storage system can set meta-information to record the amount of garbage data in the blob file. The key-value storage system can count the amount of garbage data in multiple block files through the pre-set meta-information.

[0081] A possible form of recording meta-information is as follows:

[0082] blob->garbageSize. This meta-information can be added to each blob file to record the garbage data statistics of this file.

[0083] In some embodiments, the key-value storage system can periodically or aperiodically count the amount of garbage data in multiple block files for garbage collection.

[0084] Exemplarily, the key-value storage system can count the amount of garbage data in multiple block files every once in a while for garbage collection.

[0085] In some other embodiments, in response to a garbage cleaning instruction, the key-value storage system counts the amount of garbage data in multiple block files for garbage collection.

[0086] Exemplarily, an operation and maintenance personnel can send a garbage cleaning instruction to the key-value storage system. The key-value storage system responds to this garbage cleaning instruction, obtains the amount of garbage data in multiple block files in the key-value storage system for garbage collection.

[0087] In a design, when the key-value storage system performs a background Compaction operation, the garbage data statistics of the blob file are continuously updated, that is, the key-value storage system updates the amount of garbage data in the blob file in the key-value storage system. To further improve the accuracy of garbage collection, the key-value storage system counts the latest amount of garbage data in multiple block files.

[0088] For example, the amount of garbage data of a certain blob file in the key-value storage system is 1GB at the first moment and updated to 3GB at the second moment. If the key-value storage system counts the amount of garbage data of this blob file at the second moment, the counted amount of garbage data is 3GB.

[0089] S102. The key-value storage system determines whether the amount of garbage data in each of the multiple block files meets the recycling condition.

[0090] It should be noted that the embodiments of the present application do not limit the specific recycling conditions. For example, the recycling conditions may include: the amount of garbage data is greater than or equal to the target value, or the block files with the top N garbage data amounts among multiple block files when sorted in descending order; where N is an integer greater than 0.

[0091] In some embodiments, the key-value storage system sorts multiple block files in descending order of the amount of garbage data to obtain a sorting result. If a block file is any one of the top N block files in the sorting result, the key-value storage system determines that the amount of garbage data in the block file meets the preset conditions; if a block file is not any one of the top N block files in the sorting result, the key-value storage system determines that the amount of garbage data in the block file does not meet the preset conditions.

[0092] Exemplarily, the key-value storage system is deployed with a GC module. The key-value storage system can sort multiple blob files by the amount of garbage data through the GC module, and recycle the blob files in the front positions in descending order of the amount of garbage data. For example, clean the garbage data of the top 10% of the blob files.

[0093] For another example, the key-value storage system is deployed with a GC module. The key-value storage system can sort multiple blob files by the amount of garbage data through the GC module, and recycle the top 5 blob files in descending order of the amount of garbage data.

[0094] In some other embodiments, the key-value storage system respectively compares the amount of garbage data of each block file in multiple block files with the target value. If the amount of garbage data of a block file is greater than or equal to the target value, it is determined that the amount of garbage data in the block file meets the recycling conditions; if the amount of garbage data of a block file is less than the target value, it is determined that the amount of garbage data in the block file does not meet the recycling conditions.

[0095] Exemplarily, the target value can be preset by the key-value storage system. For example, the target value can be 1GB. When performing garbage cleaning, the key-value storage system can recycle the blob files with the amount of garbage data greater than or equal to 1GB.

[0096] S103. The key-value storage system recycles the target block files that meet the recycling conditions.

[0097] In some embodiments, when the key-value storage system recycles the target block files that meet the recycling conditions, it can write the non-garbage values in the target block files into the block files outside the target block files in the key-value storage system, and delete the target block files, where the non-garbage values are the values in the target block files that have not been updated and / or deleted.

[0098] It can be understood that when recycling a target block file that meets the recycling conditions, the present application first copies the non-garbage values in the target block file to a new block file, and then deletes the target block file, which can ensure that only garbage values are deleted when cleaning up garbage in the target block file, while non-garbage values are retained.

[0099] As Figure 5 shown, the key-value storage system includes multiple blob files (such as Blob1, Blob2,...). Assuming that the blob file before the marked position is the target block file, the key-value storage system then recycles the blob file before the marked position. Specifically, for any blob file (such as Blob1) before the marked position, the key-value storage system can create a new blob file (denoted as newBlob). Further, the key-value storage system writes the non-garbage values in Blob1 into newBlob, and then deletes Blob1.

[0100] It can be understood that in the above garbage collection method, the key-value storage system performs garbage cleaning based on the garbage data volume of the blob file, and recycles the blob files whose garbage data volume meets the conditions. Compared with the related art that performs garbage cleaning based on the write time of the blob file, the garbage cleaning strategy provided by the embodiments of the present application is more in line with the actual situation of the blob file, which can improve the accuracy of garbage collection in the key-value storage system.

[0101] In some other embodiments, for a blob file without garbage data, the key-value storage system will not perform garbage cleaning on it, and thus no additional write amplification will be generated.

[0102] Exemplarily, in a pure write scenario (such as only data write operations in the key-value storage system, without read or other types of operations), the key-value storage system usually does not generate garbage data. In this case, the GC module in the key-value storage system will not rewrite any blob file data (that is, will not write the data in the old blob file into the new blob file), so no additional write amplification will be generated.

[0103] In one design, when multiple blob files are all the blob files in the key-value storage system, in order to improve the flexibility of the key-value storage system GC, as Figure 6 shown, after the above S103, the garbage collection method provided by the embodiments of the present application may further include:

[0104] S201. The key-value storage system determines the hot block files in the key-value storage system whose access frequency is greater than or equal to a preset frequency.

[0105] In a possible design, after the target block file is recycled, the key-value storage system can count the access frequencies of each blob file within a preset time period, and determine the blob files with access frequencies greater than or equal to the preset frequency as hot blob files.

[0106] Exemplarily, the access spectrum of the data in some blob files by the user is significantly higher than that of the data in other blob files. For example, the number of accesses within a week is greater than 7 times. The key-value storage system can determine this data as hot data, and the blob files where this data is located are hot blob files.

[0107] S202. The key-value storage system recycles the hot block files that meet the recycling conditions.

[0108] It should be noted that the specific implementation method of this step can refer to the above S102. The difference is that this step is for the recycling of hot blob files, that is, after the target block file is recycled, the key-value storage system can also perform garbage cleaning on the hot blob files again.

[0109] In some embodiments, when the key-value storage system executes the GC policy on all blob files, some hot blob files may not meet the criteria of the target block file. For example, all the blob files of the key-value storage system include blob file 1 - blob file 10, where blob files 7 - 10 are hot blob files. After the key-value storage system sorts blob files 1 - 10 in descending order of the amount of garbage data, blob files 7 - 10 are ranked in the last three positions, and at this time, the key-value storage system only recycles the first two ranked blob files and will not recycle blob files 7 - 10. However, after the key-value storage system finishes recycling the first two ranked blob files, it can also sort blob files 7 - 10 and perform the GC policy.

[0110] It can be understood that after the target block file is recycled, the key-value storage system performs garbage cleaning on the hot blob files again, and timely cleaning of hot garbage can be achieved by sorting the hot garbage data.

[0111] In another design, in order to reduce resource consumption during garbage collection and relieve the pressure on the key-value storage system, the multiple blob files in S101 can be hot blob files in the key-value storage system whose access frequency is greater than or equal to a preset frequency. In this case, the key-value storage system only needs to execute the GC policy for the hot blob files in the key-value storage system, without having to execute the GC policy for all blob files in the key-value storage system. The number of blob files targeted is reduced, thereby reducing resource consumption during garbage collection and relieving the pressure on the key-value storage system.

[0112] Exemplarily, the access spectrum of the data in some blob files by the user is significantly higher than that of the data in other blob files, such as the number of accesses within a week being greater than 7 times. The key-value storage system can determine this data as hot data, and the blob files where this data is located are hot blob files. The key-value storage system can only execute the above garbage collection process of S101-S103 for these hot blob files.

[0113] In one design, the SSTable file in the key-value storage system is used to store the key of the data and the address information of the value corresponding to the key. The address information is used to reflect the blob file where the value is located. To ensure the accuracy of garbage collection, the key-value storage system can continuously update the amount of garbage data in the blob file. If a certain key is updated or deleted, the key-value storage system will update the amount of garbage data in the blob file where the value corresponding to the key is located.

[0114] Exemplarily, if the first key is updated or deleted, the key-value storage system updates the amount of garbage data in the first block file to the sum of the size of the value corresponding to the first key and the original amount of garbage data in the first block file; where the first block file is the block file indicated by the address information of the value corresponding to the first key.

[0115] In some embodiments, the key-value storage system can store the identifier of the first blob file and the size of the value corresponding to the first key in the meta-information of the SSTable file where the first key is located. After the key-value storage system completes the compaction of all SSTable files, the key-value storage system reads the meta-information of the SSTable file where the first key is located and updates the amount of garbage data in the first blob file to the sum of the size of the value corresponding to the first key and the original amount of garbage data in the first blob file.

[0116] In practical applications, the key-value storage system can be configured such that the oldest blob file number is no longer recorded in the metadata of the SSTable file. Instead, all blob file information related to the SSTable file is stored, and the garbage data statistics of the current file are also recorded in the blob file. A possible form of recording the metadata is as follows:

[0117] sst-><blobs_meta>

[0118] blob->garbageSize.

[0119] Combining the above two pieces of information, a new metadata <sst, blobs> can be obtained, where sst corresponds to multiple blobs, and one blob corresponds to a garbage data volume.

[0120] In this way, after the key-value storage system completes the compaction of all SSTable files, the key-value storage system can read the metadata <sst, blobs> and update the garbage data volume of the related blob files according to the content recorded in the metadata <sst, blobs>.

[0121] It can be understood that updating the garbage data volume of the blob file together after the compaction is completed can reduce resource consumption.

[0122] In some other embodiments, during the process of the key-value storage system performing compaction on the SSTable file where the first key is located, the key-value storage system updates the garbage data volume of the first blob file to the sum of the size of the value corresponding to the first key and the original garbage data volume of the first blob file.

[0123] It can be understood that when performing compaction, when encountering keys to be updated or deleted, the garbage data volume of the corresponding blob file is updated, so that the garbage data volume of the blob file can be updated in a timely manner.

[0124] This application also provides a garbage collection device applied to a key-value storage system. In the key-value storage system, the keys and values of the data are stored in different files respectively, where the values of the data are stored in the form of block files in the key-value storage system; as Figure 7 shown, the device includes:

[0125] An acquisition module, configured to count the garbage data volume of multiple block files, and the garbage data volume of the block file is used to reflect the size of the deleted and / or updated values in the block file.

[0126] A judgment module, configured to judge whether the amount of garbage data in each of multiple block files meets the recycling condition; the recycling condition is that the amount of garbage data is greater than or equal to a target value, or the block files among the multiple block files whose garbage data amounts are sorted in descending order and ranked among the top N block files; where N is an integer greater than 0.

[0127] A processing module, configured to recycle the target block files that meet the recycling condition.

[0128] In a possible implementation manner, the judgment module is specifically configured to: sort the multiple block files in descending order of the amount of garbage data to obtain a sorting result; if a block file is any one of the top N block files in the sorting result, it is determined that the amount of garbage data in the block file meets the preset condition; if a block file is not any one of the top N block files in the sorting result, it is determined that the amount of garbage data in the block file does not meet the preset condition.

[0129] In a possible implementation manner, the judgment module is specifically configured to: respectively compare the amount of garbage data in each of the multiple block files with the target value. If the amount of garbage data in a block file is greater than or equal to the target value, it is determined that the amount of garbage data in the block file meets the recycling condition; if the amount of garbage data in a block file is less than the target value, it is determined that the amount of garbage data in the block file does not meet the recycling condition.

[0130] In a possible implementation manner, the multiple block files are all block files in a key-value storage system. After recycling the target block files that meet the recycling condition, the processing module is further configured to: determine the hot block files in the key-value storage system whose access frequency is greater than or equal to a preset frequency; recycle the hot block files that meet the recycling condition.

[0131] In a possible implementation manner, the multiple block files are hot block files in a key-value storage system whose access frequency is greater than or equal to a preset frequency.

[0132] In a possible implementation manner, the processing module is specifically configured to: write the non-garbage values in the target block file into a block file other than the target block file in the key-value storage system, and delete the target block file; the non-garbage values are the values in the target block file that have not been updated and / or deleted.

[0133] In a possible implementation manner, the processing module is further configured to: update the amount of garbage data of the block files in the key-value storage system; the acquisition module is specifically configured to: count the latest amount of garbage data of the multiple block files.

[0134] In a possible implementation, the persistent storage table file in the key-value storage system is used to store the keys of the data and the address information of the values corresponding to the keys, and the address information is used to reflect the block file where the value is located; the processing module is specifically configured to: if the first key is updated or deleted, update the garbage data volume of the first block file to the sum of the size of the value corresponding to the first key and the original garbage data volume of the first block file; wherein, the first block file is the block file indicated by the address information of the value corresponding to the first key.

[0135] In a possible implementation, the processing module is specifically configured to: store the identifier of the first block file and the size of the value corresponding to the first key into the meta-information of the persistent storage table file where the first key is located;

[0136] After the key-value storage system completes the merge of all persistent storage table files, read the meta-information of the persistent storage table file where the first key is located, and update the garbage data volume of the first block file to the sum of the size of the value corresponding to the first key and the original garbage data volume of the first block file.

[0137] In a possible implementation, the processing module is specifically configured to: during the process of the key-value storage system merging the persistent storage table file where the first key is located, update the garbage data volume of the first block file to the sum of the size of the value corresponding to the first key and the original garbage data volume of the first block file.

[0138] It should be noted that the acquisition module, the judgment module, and the processing module can all be implemented by software or by hardware. Exemplarily, next, taking the acquisition module as an example, the implementation manner of the acquisition module will be introduced. Similarly, the implementation manner of the processing module can refer to the implementation manner of the acquisition module.

[0139] As an example of a software functional unit, the acquisition module may include code running on a computing instance. Among them, the computing instance may include at least one of a physical host (computing device), a virtual machine, and a container. Further, the above computing instance may be one or more. For example, the acquisition module may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers for running this code may be distributed in the same region, or may be distributed in different regions. Further, the multiple hosts / virtual machines / containers for running this code may be distributed in the same availability zone (AZ), or may be distributed in different AZs, and each AZ includes one data center or multiple geographically proximate data centers. Usually, one region may include multiple AZs.

[0140] Similarly, multiple hosts / virtual machines / containers used to run the code can be distributed within the same virtual private cloud (VPC) or across multiple VPCs. Usually, one VPC is set up within one region. For cross-region communication between two VPCs within the same region and between VPCs in different regions, a communication gateway needs to be set up within each VPC, and the interconnection between VPCs is achieved through the communication gateway.

[0141] As an example of a hardware functional unit, the acquisition module may include at least one computing device, such as a server. Alternatively, the acquisition module may also be a device implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). Among them, the above PLD may be implemented by a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.

[0142] The multiple computing devices included in the acquisition module can be distributed in the same region or in different regions. The multiple computing devices included in the acquisition module can be distributed in the same availability zone (AZ) or in different AZs. Similarly, the multiple computing devices included in the acquisition module can be distributed within the same VPC or across multiple VPCs. Among them, the multiple computing devices may be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.

[0143] It should be noted that in other embodiments, the acquisition module can be used to execute any step in the garbage collection method, and the processing module can be used to execute any step in the garbage collection method. The steps to be implemented by the acquisition module and the processing module can be specified as needed, and the full function of the garbage collection device is achieved by separately implementing different steps in the garbage collection method through the acquisition module and the processing module.

[0144] This application also provides a garbage collection system, as Figure 8 shown, including:

[0145] A data storage device for providing data storage services, where the keys and values of the data in the data storage device are stored in different files, and the blob files in the data storage device are used to store the values of the data; A garbage collection device for: counting the amount of garbage data in multiple block files, where the amount of garbage data in a block file is used to reflect the size of the deleted and / or updated values in the block file; determining whether the amount of garbage data in each of the multiple block files meets the recycling condition; The recycling condition is: the amount of garbage data is greater than or equal to the target value, or the block files ranked in the top N in the order of the amount of garbage data in the multiple block files from largest to smallest; where N is an integer greater than 0; Recycling the target block file that meets the recycling condition.

[0146] Both the data storage device and the garbage collection device can be implemented by software or by hardware. Exemplarily, the implementation method of the data storage device will be introduced next. Similarly, the implementation method of the garbage collection device can refer to the implementation method of the data storage device.

[0147] As an example of a software functional unit, the data storage device may include code running on a computing instance. Among them, the computing instance can be at least one of computing devices such as a physical host (computing device), a virtual machine, a container, etc. Further, the above computing devices can be one or more. For example, the data storage device may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers for running the application can be distributed in the same region or in different regions. The multiple hosts / virtual machines / containers for running the code can be distributed in the same AZ or in different AZs, and each AZ includes one data center or multiple geographically proximate data centers. Among them, usually one region can include multiple AZs.

[0148] Similarly, the multiple hosts / virtual machines / containers for running the code can be distributed in the same VPC or in multiple VPCs. Among them, usually one VPC is set within one region. For cross-region communication between two VPCs within the same region and between VPCs in different regions, a communication gateway needs to be set in each VPC, and the interconnection between VPCs is achieved through the communication gateway.

[0149] As an example of a hardware functional unit, the data storage device may include at least one computing device, such as a server, etc. Or, the data storage device can also be a device implemented by ASIC, or a device implemented by PLD, etc. Among them, the above PLD can be implemented by CPLD, FPGA, GAL or any combination thereof.

[0150] The multiple computing devices included in the data storage device may be distributed in the same region or in different regions. The multiple computing devices included in the data storage device may be distributed in the same availability zone (AZ) or in different AZs. Similarly, the multiple computing devices included in the data storage device may be distributed in the same virtual private cloud (VPC) or in multiple VPCs. Among them, the multiple computing devices may be any combination of computing devices such as servers, application-specific integrated circuits (ASICs), programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), and generic array logic (GALs).

[0151] This application also provides a computing device 100. As Figure 9 shown, the computing device 100 includes: a bus 102, a processor 104, a memory 106, and a communication interface 108. The processor 104, the memory 106, and the communication interface 108 communicate with each other through the bus 102. The computing device 100 may be a server or a terminal device. It should be understood that this application does not limit the number of processors and memories in the computing device 100.

[0152] The bus 102 may be a peripheral component interconnect (PCI) bus, an extended industry standard architecture (EISA) bus, or the like. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, Figure 9 only one line is shown in the figure, but it does not mean that there is only one bus or one type of bus. The bus 102 may include a path for transmitting information between various components (such as the memory 106, the processor 104, and the communication interface 108) of the computing device 100.

[0153] The processor 104 may include any one or more of processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).

[0154] The memory 106 may include volatile memory, such as random access memory (RAM). The processor 104 may also include non-volatile memory, such as read-only memory (ROM), flash memory, a hard disk drive (HDD), or a solid state drive (SSD).

[0155] The memory 106 stores executable program code, and the processor 104 executes the executable program code to implement the functions of the aforementioned acquisition module, judgment module, and processing module respectively, thereby implementing the garbage collection method. That is, the memory 106 stores instructions for executing the garbage collection method.

[0156] Alternatively, the memory 106 stores executable code, and the processor 104 executes the executable code to implement the functions of the aforementioned data storage device and garbage collection device respectively, thereby implementing the garbage collection method. That is, the memory 106 stores instructions for executing the garbage collection method.

[0157] The communication interface 108 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 100 and other devices or communication networks.

[0158] The embodiments of the present application also provide a computing device cluster. The computing device cluster includes at least one computing device. The computing device may be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device may also be a terminal device such as a desktop computer, a laptop computer, or a smart phone.

[0159] As Figure 10 shown, the computing device cluster includes at least one computing device 100. The memory 106 in one or more of the computing devices 100 in the computing device cluster may store the same instructions for executing the garbage collection method.

[0160] In some possible implementation manners, the memory 106 in one or more of the computing devices 100 in the computing device cluster may also store partial instructions for executing the garbage collection method respectively. In other words, the combination of one or more computing devices 100 may jointly execute the instructions for executing the garbage collection method.

[0161] It should be noted that the memories 106 in different computing devices 100 in the computing device cluster may store different instructions, respectively for executing partial functions of the YY device. That is, the instructions stored in the memories 106 in different computing devices 100 can implement the functions of one or more of the acquisition module, the judgment module, and the processing module.

[0162] In some possible implementation manners, one or more computing devices in the computing device cluster may be connected through a network. Among them, the network may be a wide area network or a local area network, etc. Figure 11 A possible implementation manner is shown. As Figure 11 shown, three computing devices 100A, 100B, and 100C are connected through a network. Specifically, they are connected to the network through the communication interfaces in each computing device. In this type of possible implementation manner, the memory 106 in the computing device 100A stores instructions for executing the function of the acquisition module. At the same time, the memory 106 in the computing device 100B stores instructions for executing the judgment module, and the memory 106 in the computing device 100C stores instructions for executing the function of the processing module.

[0163] Figure 11 The connection manner between the computing device clusters shown may be considered that since the garbage collection method provided in this application needs to process a large amount of data, it is considered to hand over the functions implemented by the acquisition module, the judgment module, and the processing module to the computing device 100B for execution.

[0164] It should be understood that Figure 11 the function of the computing device 100A shown in [[ ]] can also be completed by multiple computing devices 100. Similarly, the function of the computing device 100B can also be completed by multiple computing devices 100, and the function of the computing device 100C can also be completed by multiple computing devices 100.

[0165] The embodiments of this application also provide another computing device cluster. The connection relationship between the computing devices in this computing device cluster can be similarly referred to Figure 10 and Figure 11 the connection manner of the described computing device cluster. The difference is that the memories 106 in one or more computing devices 100 in this computing device cluster may store the same instructions for executing the garbage collection method.

[0166] In some possible implementation manners, the memories 106 in one or more computing devices 100 in this computing device cluster may also respectively store partial instructions for executing the garbage collection method. In other words, a combination of one or more computing devices 100 can jointly execute the instructions for executing the garbage collection method.

[0167] It should be noted that the memories 106 in different computing devices 100 in the computing device cluster may store different instructions for performing some functions of the garbage collection system. That is, the instructions stored in the memories 106 in different computing devices 100 can implement the functions of one or more of the data storage device and the garbage collection device.

[0168] An embodiment of the present application further provides a computer program product containing instructions. The computer program product may be software or a program product containing instructions that can run on a computing device or be stored in any available medium. When the computer program product runs on at least one computing device, at least one computing device is caused to execute the garbage collection method of the embodiment of the present application.

[0169] An embodiment of the present application further provides a computer-readable storage medium. The computer-readable storage medium may be any available medium that a computing device can store or a data storage device such as a data center containing one or more available media. The available medium may be a magnetic medium (e.g., a floppy disk, a hard disk, a magnetic tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive), etc. The computer-readable storage medium includes instructions that instruct a computing device to execute the garbage collection method of the embodiment of the present application.

[0170] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the protection scope of the technical solutions of the embodiments of the present invention.

Claims

1. A garbage collection method, characterized in that: Applied to a key-value storage system, in which the key and value of data are stored in different files respectively, wherein the value of the data is stored in the form of a block file in the key-value storage system; the method comprises: The key-value storage system counts the amount of garbage data in a plurality of block files, where the amount of garbage data in the block files is used to reflect the size of the value deleted and / or updated in the block files; The key-value storage system determines whether the amount of garbage data of each block file among the multiple block files meets a recycling condition; the recycling condition is: the amount of garbage data is greater than or equal to a target value, or the amount of garbage data of the multiple block files is sorted in descending order and is ranked in the first N block files; wherein N is an integer greater than 0; The key-value storage system recycles the target block file that meets the recycle condition.

2. The method according to claim 1, characterized in that The key-value storage system determines whether the amount of garbage data in each block file among the plurality of block files meets a recycling condition, including: The key-value storage system sorts the multiple block files in descending order of the amount of garbage data to obtain a sorting result; If the block file is any one of the top N block files in the sorting result, the key-value storage system determines that the amount of junk data of the block file meets a preset condition; If the block file is not any of the top N block files in the sorting result, the key-value storage system determines that the amount of junk data of the block file does not meet a preset condition.

3. The method according to claim 1, characterized in that The key-value storage system determines whether the amount of garbage data in each block file among the plurality of block files meets a recycling condition, including: The key-value storage system compares the amount of garbage data of each block file among the multiple block files with the target value respectively. If the amount of garbage data of the block file is greater than or equal to the target value, it is determined that the amount of garbage data of the block file meets the recycling condition; if the amount of garbage data of the block file is less than the target value, it is determined that the amount of garbage data of the block file does not meet the recycling condition.

4. The method according to claim 1, characterized in that The multiple block files are all block files in the key-value storage system. After the key-value storage system recycles the target block files that meet the recycling condition, the method further includes: Determine a hotspot block file in the key-value storage system whose access frequency is greater than or equal to a preset frequency; Reclaim the hotspot block files that meet the recycling condition.

5. The method according to claim 1, characterized in that The plurality of block files are hotspot block files in the key-value storage system whose access frequency is greater than or equal to a preset frequency.

6. The method according to any one of claims 1 to 5, characterized in that: The key-value storage system reclaims target block files that meet the reclaiming conditions, including: The key-value storage system writes the non-garbage values ​​in the target block file to block files outside the target block file in the key-value storage system, and deletes the target block file; the non-garbage values ​​are values ​​that have not been updated and / or deleted in the target block file.

7. The method according to any one of claims 1 to 6, characterized in that: The method further comprises: The key-value storage system updates the amount of garbage data of the block files in the key-value storage system; The key-value storage system counts the amount of garbage data of multiple block files, including: The key-value storage system counts the latest garbage data amounts of multiple block files.

8. The method according to claim 7, characterized in that The persistent storage table file in the key-value storage system is used to store the key of the data and the address information of the value corresponding to the key, and the address information is used to reflect the block file where the value is located; The key-value storage system updates the amount of garbage data of a block file in the key-value storage system, including: If the first key is updated or deleted, the key-value storage system updates the amount of garbage data in the first block file to the sum of the size of the value corresponding to the first key and the original amount of garbage data in the first block file; wherein the first block file is the block file indicated by the address information of the value corresponding to the first key.

9. The method according to claim 8, characterized in that The key-value storage system updates the amount of garbage data in the first block of files to the sum of the value corresponding to the first key and the original amount of garbage data in the first block of files, including: storing the identifier of the first block file and the size of the value corresponding to the first key in the meta information of the persistent storage table file where the first key is located; After the key-value storage system completes merging of all persistent storage table files, the meta information of the persistent storage table file where the first key is located is read, and the amount of garbage data in the first block file is updated to be the sum of the size of the value corresponding to the first key and the original amount of garbage data in the first block file.

10. The method according to claim 8, characterized in that The key-value storage system updates the amount of garbage data in the first block of files to the sum of the value corresponding to the first key and the original amount of garbage data in the first block of files, including: When the key-value storage system merges the persistent storage table files where the first key is located, the amount of garbage data in the first block file is updated to the sum of the value corresponding to the first key and the original amount of garbage data in the first block file.

11. A garbage collection device, characterized in that: Applied to a key-value storage system, in which the key and value of data are stored in different files respectively, wherein the value of the data is stored in the form of a block file in the key-value storage system; the device comprises: An acquisition module, used for counting the amount of junk data of multiple block files, where the amount of junk data of the block files is used to reflect the size of the value deleted and / or updated in the block files; a judgment module, configured to judge whether the amount of garbage data of each block file among the plurality of block files meets a recycling condition; the recycling condition is: the amount of garbage data is greater than or equal to a target value, or the amount of garbage data of the plurality of block files is sorted in descending order and is ranked in the first N block files; wherein N is an integer greater than 0; The processing module is used to recycle the target block files that meet the recycling conditions.

12. The device according to claim 11, characterized in that The judgment module is specifically used for: Sorting the plurality of block files in descending order of the amount of garbage data to obtain a sorting result; If the block file is any one of the top N block files in the sorting result, it is determined that the amount of junk data of the block file meets the preset condition; If the block file is not any of the top N block files in the sorting result, it is determined that the amount of junk data of the block file does not meet the preset condition.

13. The device according to claim 11, characterized in that The judgment module is specifically used for: The amount of garbage data of each block file among the multiple block files is compared with the target value respectively. If the amount of garbage data of the block file is greater than or equal to the target value, it is determined that the amount of garbage data of the block file meets the recycling condition. If the amount of garbage data of the block file is less than the target value, it is determined that the amount of garbage data of the block file does not meet the recycling condition.

14. The device according to claim 11, characterized in that The multiple block files are all block files in the key-value storage system. After reclaiming the target block files that meet the reclaiming condition, the processing module is further used to: Determine a hotspot block file in the key-value storage system whose access frequency is greater than or equal to a preset frequency; Reclaim the hotspot block files that meet the recycling condition.

15. The device according to claim 11, characterized in that The plurality of block files are hotspot block files in the key-value storage system whose access frequency is greater than or equal to a preset frequency.

16. The device according to any one of claims 11 to 15, characterized in that: The processing module is specifically used for: Writing non-garbage values ​​in the target block file to block files other than the target block file in the key-value storage system, and deleting the target block file; the non-garbage values ​​are values ​​in the target block file that have not been updated and / or deleted.

17. The device according to any one of claims 11 to 16, characterized in that: The processing module is also used for: Updating the amount of garbage data in the block file in the key-value storage system; The acquisition module is specifically used for: Count the latest amount of garbage data in multiple block files.

18. The device according to claim 17, characterized in that The persistent storage table file in the key-value storage system is used to store the key of the data and the address information of the value corresponding to the key, and the address information is used to reflect the block file where the value is located; the processing module is specifically used to: If the first key is updated or deleted, the amount of garbage data in the first block file is updated to the sum of the value corresponding to the first key and the original amount of garbage data in the first block file; wherein the first block file is the block file indicated by the address information of the value corresponding to the first key.

19. The device according to claim 18, characterized in that The processing module is specifically used for: storing the identifier of the first block file and the size of the value corresponding to the first key in the meta information of the persistent storage table file where the first key is located; After the key-value storage system completes merging of all persistent storage table files, the meta information of the persistent storage table file where the first key is located is read, and the amount of garbage data in the first block file is updated to be the sum of the size of the value corresponding to the first key and the original amount of garbage data in the first block file.

20. The device according to claim 18, characterized in that The processing module is specifically used for: When the key-value storage system merges the persistent storage table files where the first key is located, the amount of garbage data in the first block file is updated to the sum of the value corresponding to the first key and the original amount of garbage data in the first block file.

21. A computing device cluster, characterized in that: comprising at least one computing device, each computing device comprising a processor and a memory; The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster executes the garbage collection method according to any one of claims 1 to 10.

22. A computer program product comprising instructions, characterized in that When the instruction is executed by a computing device cluster, the computing device cluster executes the garbage collection method according to any one of claims 1 to 10.

23. A computer-readable storage medium, characterized in that: The method comprises computer program instructions. When the computer program instructions are executed by a computing device cluster, the computing device cluster executes the garbage collection method according to any one of claims 1 to 10.