Techniques for garbage collection for zones in file system
By assigning sequence numbers to file system areas and combining them with garbage rate and dwell time, candidate areas are intelligently selected for garbage collection, solving the problems of write amplification and data mixing in existing technologies, and achieving more efficient resource utilization and data separation.
Patent Information
- Application Number
- CN202411679486.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-04-11
- Filing Date
- 2024-11-22
- Publication Date
- 2025-10-21
AI Technical Summary
Existing file system garbage collection algorithms fail to effectively distinguish data hotness, resulting in write amplification and a mixture of hot and cold data. Furthermore, they fail to intelligently select GC candidate regions, increasing write amplification and resource waste.
By assigning a sequence number to each region and combining the garbage rate and residence time, candidate regions are intelligently selected for garbage collection. Safe and critical GC processes are adopted to optimize candidate region selection under different capacity thresholds, avoiding GC in hot data regions and improving data separation and resource utilization efficiency.
It effectively avoids write amplification, improves data separation, reduces write amplification, and enhances resource utilization and GC compression efficiency.
Smart Images

Figure CN120821705A_ABST
Abstract
Description
Background Art
[0001] The described aspects relate to file systems and, more particularly, to mechanisms for performing garbage collection on extents of a file system.
[0002] Some file systems (such as ZenFS) store data in zones on raw partitioned block devices. A zone can be a logical construct for storing data as well as metadata. The metadata can include attributes associated with the zone, which can include indicators as to whether certain data in the zone is valid or invalid (e.g., overwriting existing data, in a different zone, or otherwise no longer relevant). Such file systems can store multiple files in a single zone by using an extent allocation scheme. A file can be composed of one or more segments, and all the segments that make up the file can be stored in the same zone (or different zones) of the device. When all file segments in a zone are invalid, the zone can be reset and then reused to store new file segments. In some such file systems, when the storage capacity of the drive drops below a certain threshold, the file system can initiate garbage collection (GC) of the zone. During GC, the file system can relocate valid data from one or more zones to other (e.g., newer) zones, and can proceed to erase one or more zones once only invalid data remains.
[0003] When selecting extents for GC, the file system can prioritize them based on a garbage rate metric, which can be a measure of invalid data compared to the total data in the extent. Current GC algorithms simply scan all extents and select the best candidate with the greatest invalid data. This allows the file system to free up the most space by performing GC on the extent with the greatest amount of invalid data (and / or potentially overwriting the least amount of valid data). However, this algorithm may have some drawbacks. First, current GC algorithms may not consider data heat, where heat can refer to how recently data has been written or rewritten. For example, because the algorithm does not consider data heat or coldness, it may always end up selecting extents with hot data. In many workloads, extents containing hot data are frequently overwritten by the host. Consequently, the host itself may continue to invalidate data, ultimately causing the entire extent to be erased rather than undergoing GC. Second, current GC algorithms may result in increased write amplification. For example, without any intelligence, the file system may select an extent candidate for GC that has already been scheduled for complete erasure by the host, resulting in increased write amplification. Third, the current GC algorithm may lead to a lack of separation between hot and cold data - for example, because the algorithm selects region candidates based only on the highest garbage rate, over time, regions may become a mixture of hot and cold data. Therefore, garbage collection of such regions may lead to increased write amplification. Summary of the Invention
[0004] The following presents a simplified summary of one or more embodiments in order to provide a basic understanding of such embodiments. This summary is not an extensive overview of all contemplated embodiments and is not intended to identify key or critical elements of all embodiments, nor to delineate the scope of any or all embodiments. Its sole purpose is to present some concepts of one or more embodiments in a simplified form as a prelude to the more detailed description that is presented later.
[0005] In an example, a computer-implemented method for performing garbage collection in a file system having a plurality of data areas is provided, the method comprising calculating, for each of the plurality of areas in the file system, a garbage rate associated with an amount of invalid data in the area, determining one or more candidate areas for garbage collection in the plurality of areas based on the garbage rate and a sequence number assigned to each area, and performing garbage collection on the one or more candidate areas in the file system.
[0006] In another example, an apparatus for performing garbage collection in a file system having multiple data areas is provided, the apparatus comprising one or more processors and one or more non-transitory memories having instructions thereon. The instructions, when executed by the one or more processors, cause the one or more processors to calculate, for each of the multiple areas in the file system, a garbage rate associated with an amount of invalid data in the area, determine one or more candidate areas for garbage collection in the multiple areas based on the garbage rate and a sequence number assigned to each area, and perform garbage collection on the one or more candidate areas in the file system.
[0007] In another example, one or more non-transitory computer-readable storage media storing instructions are provided that, when executed by one or more processors, cause the one or more processors to perform a method for performing garbage collection in a file system having a plurality of data areas. The method includes calculating, for each of a plurality of areas in the file system, a garbage rate associated with an amount of invalid data in the area, determining one or more candidate areas for garbage collection from the plurality of areas based on the garbage rate and a sequence number assigned to each area, and performing garbage collection on the one or more candidate areas in the file system.
[0008] To accomplish the foregoing and related ends, one or more embodiments comprise the features hereinafter fully described and particularly pointed out in the claims. The following description and the accompanying drawings set forth in detail certain illustrative features of one or more embodiments. However, these features are indicative of but a few of the various ways in which the principles of various embodiments may be employed, and this description is intended to include all such embodiments and their equivalents. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] Figure 1is a schematic diagram of an example of a system for performing garbage collection (GC) in a file system according to examples described herein.
[0010] Figure 2 is a flow chart of an example of a method for performing GC in a file system according to examples described herein.
[0011] Figures 3A-3D is a flow chart of a specific example of a method for determining candidate regions for GC in a file system according to aspects described herein.
[0012] Figure 4 is a schematic diagram of an example of an apparatus for performing the functions described herein. DETAILED DESCRIPTION
[0013] The detailed description set forth below in conjunction with the accompanying drawings is intended as a description of various configurations and is not intended to represent the only configuration in which the concepts described herein may be practiced. The detailed description includes specific details for the purpose of providing a thorough understanding of the various concepts. However, it will be apparent to those skilled in the art that these concepts may be practiced without these specific details. In some instances, well-known components are shown in block diagram form to avoid obscuring such concepts.
[0014] The present disclosure describes various examples related to performing garbage collection (GC) on one or more zones of a file system, including selecting one or more candidate zones for GC based on a garbage rate associated with the amount of invalid data in the zone and a sequence number assigned to the zone. For example, a sequence number can be assigned to each zone to help determine candidate zones for GC. In one example, only zones that are not within the residence time of the most recently allocated zone can be considered for GC. The residence time can be defined by the sequence number such that only zones with a sequence number that is at least the number of residence times away from the last allocated sequence number can be considered candidates for GC. Thus, the residence time can represent a range of sequence numbers starting from the sequence number of the most recently allocated zone to be avoided during GC. This can avoid GC of zones containing hot data (e.g., data that has been recently written or rewritten), which can avoid write amplification because, in many cases, the host application can consistently overwrite the hottest data. In another example, a device capacity threshold can be defined that affects how candidate selection is performed.
[0015] In one example, in a case where a device executing a file system has a capacity less than a first threshold, a safe GC process may be used to select candidate areas for GC. In the safe GC process, multiple areas may be selected as candidate areas based on garbage rate and / or compliance with residence time. Multiple candidate areas may be refined based on other dimensions. One dimension may include whether the garbage rate of a given candidate area is within the garbage rate gap of a candidate area with the highest garbage rate, which may ensure that only candidate areas with a certain amount of garbage data relative to the highest amount of garbage data may be GCed. Another dimension may include whether the sequence number of a given candidate area is within the sequence number gap of a candidate area with the lowest sequence number, which may ensure that only areas with similar cold data are considered for GC. In another example, in a case where a device executing a file system has a capacity less than a second threshold, which is less than the first threshold, a critical GC process may be used to select candidate areas for GC. In the critical GC process, an area with the highest garbage rate and also meeting the residence time may be selected as a candidate area for GC.
[0016] According to various aspects described herein, a modified GC process can use the sequence number of an extent to distinguish between hot and cold data, different thresholds for distinguishing system actions when triggered, dwell times, and the like. In this regard and as described herein, using sequence numbers can allow for a more accurate assessment of the hotness and coldness of data beyond simply relying on lifecycle hints that the file system may provide in metadata. For example, when comparing two extents, if the extent with the lower sequence number still contains valid data compared to the extent with the higher sequence number, the system can infer that the data in the extent with the lower sequence number is colder. As previously described, the hotness and coldness of data can refer to the time at which data was appended to the file system (e.g., hotter data is appended earlier than colder data). Additionally, in some examples, implementing thresholds for determining GC processes (e.g., critical and safe) can provide a mechanism to identify the urgency of performing GC. For example, during a critical GC process, the file system can ensure that the drive has enough space to accommodate new writes by selecting the extent with the highest garbage rate, and a critical GC process may require less processing than a safe GC process and, therefore, can be more efficient in write-heavy scenarios. During safe GC, the file system can exhibit enhanced intelligence when selecting GC candidates, as described above and further herein, which can allow for improved GC compaction, improved separation of hot and cold data, reduced write amplification, and the like. Using dwell time can be used to prevent selecting extents from the hot stream, thereby providing the host with sufficient time to invalidate data, which can help improve write amplification. In other examples described herein, the file system can maintain separate GC extent streams, which can also help separate hot and cold data.
[0017] As used herein, a processor, at least one processor and / or one or more processors, configured individually or in combination to perform or be operable to perform multiple actions means including at least two different processors that can perform different, overlapping or non-overlapping subsets of multiple actions, or a single processor that can perform all actions in multiple actions. In a non-limiting example where multiple processors can perform different actions in combination, the description of a processor, at least one processor and / or one or more processors that are configured or can be operated to perform actions X, Y and Z may include at least a first processor that is configured or can be operated to perform a first subset of X, Y and Z (e.g., perform X) and at least a second processor that is configured or can be operated to perform a second subset of X, Y and Z (e.g., perform Y and Z). Alternatively, the first processor, the second processor and the third processor can be configured or operable to perform a corresponding one of actions X, Y and Z, respectively. It should be understood that any combination of one or more processors can be configured or operable to perform any one action or any combination of multiple actions.
[0018] As used herein, a memory, at least one memory, and / or one or more memories, individually or in combination, configured to store or having stored thereon instructions executable by one or more processors for performing a plurality of actions, means including at least two different memories capable of storing different, overlapping, or non-overlapping subsets of instructions for performing different, overlapping, or non-overlapping subsets of the plurality of actions, or a single memory capable of storing instructions for performing all of the plurality of actions. In one non-limiting example where one or more memories (individually or in combination) are capable of storing different subsets of instructions for performing different ones of the plurality of actions, a description of a memory, at least one memory, and / or one or more memories configured or operable to store or having stored thereon instructions for performing actions X, Y, and Z may include at least a first memory configured or operable to store or having stored thereon instructions for performing a first subset of X, Y, and Z (e.g., instructions for performing X) and at least a second memory configured or operable to store or having stored thereon instructions for performing a second subset of X, Y, and Z (e.g., instructions for performing Y and Z). Alternatively, the first memory, the second memory, and the third memory may each be configured to store or store thereon a corresponding one of a first subset of instructions for performing X, a second subset of instructions for performing Y, and a third subset of instructions for performing Z. It should be understood that any combination of one or more memories may be configured or operable to store or store thereon any one or any combination of instructions executable by one or more processors to perform any one or any combination of a plurality of actions. Furthermore, one or more processors may each be coupled to at least one of the one or more memories and configured or operable to execute instructions to perform the plurality of actions. For example, in the non-limiting example of different subsets of instructions for performing actions X, Y, and Z described above, a first processor may be coupled to a first memory storing instructions for performing action X, and at least a second processor may be coupled to at least a second memory storing instructions for performing actions Y and Z, and the first and second processors may, in combination, execute the instructions of the corresponding subsets to complete the execution of actions X, Y, and Z. Alternatively, the three processors may access one of three different memories, each storing one of the instructions for performing X, Y, or Z, and the three processors may in combination execute corresponding subsets of the instructions to complete the performance of actions X, Y, and Z. Alternatively, a single processor may execute instructions stored on a single memory or distributed across multiple memories to complete the performance of actions X, Y, and Z.
[0019] Now go to Figures 1-4, examples are depicted with reference to one or more components and one or more methods that can perform the actions or operations described herein, where components and / or actions / operations in dashed lines may be optional. Figure 2 and Figures 3A-3D The operations described in the specification are presented in a particular order and / or performed by example components, but in some examples, the ordering of the actions and the components performing the actions may vary depending on the implementation. Furthermore, in some examples, one or more of the actions, functions, and / or described components may be performed by a specially programmed processor, a processor executing specially programmed software or computer-readable media, or any other combination of hardware components and / or software components capable of performing the described actions or functions.
[0020] Figure 1 1 is a schematic diagram of an example of a system for performing GC in a file system according to aspects described herein. The system includes a device 100 (e.g., a computing device) that includes processor(s) 102 (e.g., one or more processors) and / or memory / memories 104 (e.g., one or more memories). In an example, the device 100 may include processor(s) 102 and / or memory / memories 104 configured to execute or store instructions or other parameters related to providing an operating system 106, which may execute one or more applications, services, etc. In another example, the device 100 may execute a host application 110, for example, via the operating system 106. For example, the host application 110 may include a user application that may create, modify, update, etc. data, which may be stored by the file system 120.
[0021] For example, the processor(s) 102 and the memory / memories 104 may be separate components communicatively coupled via a bus (e.g., on a motherboard or other portion of a computing device, on an integrated circuit such as a system on a chip (SoC), etc.), components integrated with one another (e.g., the processor(s) 102 may include the memory / memories 104 as an onboard component 101), etc. In other examples, the processor(s) 102 may include multiple processors 102 of multiple devices 100, the memory / memories 104 may include multiple memories 104 of multiple devices 100, etc. The memory / memories 104 may store instructions, parameters, data structures, etc. for use / execution by the processor(s) 102 to perform the functionality described herein.
[0022] Furthermore, the device 100 may include substantially any device that may have a processor(s) 102 and a memory / memories 104, such as a computer (e.g., a workstation, a server, a personal computer, etc.), a personal device (e.g., a cellular phone, such as a smartphone, a tablet, etc.), a smart device (such as a smart TV, etc.). Furthermore, in examples, various components or modules of the device 100 may be within a single device, as shown.
[0023] In an example, the file system 120 may be a file system, such as ZenFS, that stores data in zones on a raw partitioned block storage device 130. A zone may be defined as a logical space on the block storage device 130 for storing data and / or associated metadata for the zone, where the metadata may include one or more parameters for the zone, such as an indicator for each data in the zone indicating whether the data is valid or invalid. The block storage device 130 may include a hard disk drive (HDD), a solid-state drive (SSD), or substantially any form of memory (e.g., random access memory (RAM), read-only memory (ROM), tape, magnetic disk, optical disk, volatile memory, non-volatile memory, and any combination thereof). The file system 120 may store data in a zone until the zone is full, and then may continue to store data in the next or newly created zone. According to various aspects described herein, the file system 120 may assign a sequence number to a newly created zone. In an example, data being overwritten or otherwise modified may be stored in the newest zone, and previous data (possibly in a different zone) may be marked as invalid (e.g., in the zone metadata).
[0024] For example, the host application 110 may request to store data in the file system 120, and the file system 120 may optionally include an extent management module 122 for handling extent management and data storage, marking overwritten data as invalid, etc. In an example, the host application 110 may also request or cause data to be deleted from the file system 120, which may result in the data being marked as invalid. The file system 120 may operate a separate garbage collection process to delete invalid data to create space for new data (or corresponding extents). In another example, the file system 120 may optionally include a candidate determination module 124 for determining one or more candidate extents for garbage collection, and / or a garbage collection module 126 for performing garbage collection on one or more candidate extents.
[0025] Figure 2 is a flow chart of an example of a method 200 for performing GC in a file system according to aspects described herein. For example, the method 200 may be performed by the device 100 executing the file system 120.
[0026] In method 200, at act 202, a garbage rate associated with the amount of invalid data in a zone can be calculated for each of a plurality of zones in a file system. For example, zone management module 122, e.g., in conjunction with processor(s) 102, memory / memories 104, operating system 106, file system 120, etc., can calculate a garbage rate associated with the amount of invalid data in the zone for each of a plurality of zones in a file system (e.g., file system 120). For example, zone management module 122 can calculate the garbage rate periodically, or when data in the zone is modified (e.g., added or marked invalid), or as part of a startup garbage collection process, etc. In an example, zone management module 122 can calculate the garbage rate, or calculate the garbage rate based on the size of data marked invalid in the zone, the ratio of invalid data to total data (or valid data), etc. Zone management module 122 can determine the amount of invalid data in the zone based on zone metadata, which can indicate, for each data in the zone, whether the data is valid or invalid.
[0027] In method 200, at act 204, one or more candidate zones for garbage collection from a plurality of zones may be determined based on a garbage rate and a sequence number assigned to each zone. For example, candidate determination module 124, e.g., in conjunction with processor(s) 102, memory / memories 104, operating system 106, file system 120, etc., may determine one or more candidate zones for GC from a plurality of zones based on a garbage rate and a sequence number assigned to each zone. For example, candidate determination module 124 may analyze the garbage rate of each zone, which may be affected by one or more dimensions as described herein, and / or the sequence number of each zone to determine which zones are to be considered for GC. As previously described, for example, file system 120 may assign a sequence number to a zone for storing data when creating the zone (e.g., when creating each zone), such that the sequence number increments for each zone. For example, file system 120 may assign a sequence number of zero or one to the first created zone, and may then assign an increasing number to each subsequently created zone. In this regard, extents with higher sequence numbers may correspond to recently created extents, which may include "hotter" data, compared to extents with lower sequence numbers (which correspond to older extents, which may include "colder" data). Using sequence numbers may allow hot data not to be GCed, which might otherwise be overwritten by the host application 110 for GC.
[0028] When determining one or more candidate zones at act 204, optionally at act 206, a determination may be made as to whether the sequence number of the one or more candidate zones is outside a dwell time from a highest sequence number in the plurality of zones. For example, the candidate determination module 124, e.g., in conjunction with the processor(s) 102, the memory / memories 104, the operating system 106, the file system 120, etc., may determine that the sequence number of the one or more candidate zones is outside a dwell time from a highest sequence number in the plurality of zones. In an example, the dwell time may be configured in the file system 120, may be configured by the host application 110, etc. The dwell time may be configured based on a number of sequence numbers such that the candidate determination module 124 may verify that the one or more candidate zones are not within a number of sequence numbers from a sequence number of a last created zone. If a zone is within the dwell time, the candidate determination module 124 may exclude the zone from being a candidate zone.
[0029] For example, the sequence number can be a global number that is incremented each time a new zone is opened and can be part of the zone structure (and can be power safe so that it is preserved in the event of a power outage). The dwell time can represent the duration that the GC process avoids selecting a candidate zone for a specified time frame represented by the sequence number. For example, the candidate determination module 124 can decide not to consider the zone candidates of the latest n sequence numbers, where n is a positive integer. In this example, the fixed duration can be referred to as the "dwell time", which is set to n. In the example, the dwell time can be considered only for zones with live writes or hot flows, and not for GC flows, where GC flows can include data that is migrated as part of GC of the zone where the data was previously located.
[0030] When determining one or more candidate regions at act 204, optionally at act 208, the candidate regions may include multiple candidate regions, and a subset of the multiple candidate regions having the highest junk rate in the one or more candidate regions may be determined. For example, the candidate determination module 124, for example, in conjunction with the processor(s) 102, the memory / memories 104, the operating system 106, the file system 120, etc., may iterate through the multiple regions (or at least the multiple regions outside of the dwell time) to determine a subset of the multiple candidate regions in the one or more candidate regions that have the highest junk rate (e.g., in the multiple regions or at least in the multiple regions outside of the dwell time). In an example, the candidate determination module 124 may determine that the subset includes a fixed number of candidate regions and / or may sort the subset from highest junk rate to lowest junk rate. As previously described, the junk rate for each region may be calculated by the region management module 122 (e.g., at act 202), stored in the region metadata, etc.
[0031] Upon determining one or more candidate regions at act 204, at act 210, at least one candidate region having a garbage rate less than a garbage rate gap from the highest garbage rate of the highest candidate region in the subset may be optionally removed from the subset. For example, candidate determination module 124, e.g., in conjunction with processor(s) 102, memory / memories 104, operating system 106, file system 120, etc., may remove from the subset at least one candidate region having a garbage rate less than a garbage rate gap from the highest garbage rate of the highest candidate region in the subset. In this regard, for example, file system 120 or host application 110 may define the garbage rate gap as the maximum gap between the garbage rate of the candidate region with the highest garbage rate (and outside of the residence time) and the other candidate regions in the subset considered for garbage collection. Thus, the garbage rate gap may represent a range of garbage rates, starting from the highest garbage rate, to be considered during garbage collection. In an example, the garbage rate gap may be configurable by file system 120. For example, the garbage rate gap may refer to a permissible percentage difference considered when comparing GC candidate regions. Candidate regions with garbage rates exceeding this threshold may not be considered for GC. In some examples, this dimension can have the highest priority over any other dimension because it can avoid write amplification. For example, if the garbage rate gap is 10%, and the highest candidate in the subset being considered for GC has a garbage rate of 60%, then the subset of candidate regions being considered will have a garbage rate of at least 50%.
[0032] When determining one or more candidate regions at act 204, at least one candidate region having a higher sequence number than other candidate regions in the subset can optionally be removed from the subset at act 212. For example, the candidate determination module 124, e.g., in conjunction with the processor(s) 102, the memory / memories 104, the operating system 106, the file system 120, etc., can remove from the subset at least one candidate region having a higher sequence number than other candidate regions in the subset. In an example, the candidate determination module 124 can remove one or more candidate regions having the highest sequence number from the subset to generate a subset of a specific size or to ensure a specific difference in sequence numbers among the subsets.
[0033] When determining one or more candidate regions at act 204, optionally at act 214, at least one candidate region whose sequence number is greater than a sequence number gap from the candidate region with the lowest sequence number can be removed from the subset. For example, the candidate determination module 124, for example, in conjunction with the processor(s) 102, the memory / memories 104, the operating system 106, the file system 120, etc., can remove at least one candidate region from the subset whose sequence number is greater than a sequence number gap from the candidate region with the lowest sequence number. This can ensure that the candidate region is within a sequence number gap from the lowest sequence number present in the subset. For example, the sequence number gap can represent a range of sequence numbers starting from the lowest sequence number to be considered when garbage collecting. In an example, the sequence number gap can be configured by the file system 120. The sequence number gap can represent a range to be considered after identifying the top GC candidates based on their garbage rates. Candidates within this sequence number gap are considered to have similar cold data, and the candidate with the lowest sequence number can be prioritized for GC.
[0034] When determining one or more candidate zones at act 204, optionally at act 216, one or more candidate zones having the highest garbage rate and having a sequence number that is outside the residence time from the highest sequence number in the plurality of zones may be determined. For example, the candidate determination module 124, e.g., in conjunction with the processor(s) 102, the memory / memories 104, the operating system 106, the file system 120, etc., may determine one or more of the candidate zones having the highest garbage rate and having a sequence number that is outside the residence time from the highest sequence number in the plurality of zones. The highest sequence number in the plurality of zones may correspond to the last added zone. As previously described, the candidate zone(s) for GC may be outside the residence time from the last added zone. Thus, in this example, one or more candidate zones (e.g., a single zone or a group of zones) having the highest garbage rate outside the residence time may be selected for GC.
[0035] When determining one or more candidate areas at act 204, optionally at act 218, it may be determined that one or more candidate areas are open by the host application. For example, candidate determination module 124, e.g., in conjunction with processor(s) 102, memory / memories 104, operating system 106, file system 120, etc., may determine that one or more of the candidate areas are open by the host application 110. In an example, areas may be created by file system 120 for relocating data from GCed areas, and areas may also be created (e.g., opened) by host application 110 for storing or rewriting data. In an example, candidate determination module 124 may distinguish between these areas and may ignore areas created by file system 120 (rather than by or on behalf of host application 110) when determining candidate areas for GC. In one example, areas opened by file system 120 for relocating data may be considered GC streams, and GC streams may be avoided when determining candidate areas for GC. In some file systems, extent migration only considers lifecycle hints, resulting in a mix of areas containing hot and cold data. This may significantly hinder GC compaction. In this regard, using GC streams to write migrated ranges can avoid compaction issues.
[0036] In method 200, optionally at act 220, a device capacity can be compared to a threshold. For example, candidate determination module 124, e.g., in conjunction with processor(s) 102, memory / memories 104, operating system 106, file system 120, etc., can compare the device capacity (e.g., the capacity of block storage device 130) to the threshold. In an example, candidate determination module 124 can perform this comparison when determining which of acts 206, 208, 210, 212, 214, 216, and / or 218 to perform to determine one or more candidate blocks for GC. In one example, candidate determination module 124 can compare the device capacity to a first threshold (e.g., a safety threshold), and if the device capacity is less than the safety threshold but greater than a second threshold (e.g., a critical threshold), candidate determination module 124 can determine to perform certain actions to determine candidate regions, such as acts 206, 208, 210, 212, and / or 214. In another example, the candidate determination module 124 may compare the device capacity to a second threshold (e.g., a critical threshold), and if the device capacity is less than the critical threshold, the candidate determination module 124 may determine to perform certain actions to determine candidate regions, such as actions 206 and / or 216.
[0037] In method 200, at act 222, garbage collection of one or more candidate regions may be performed. For example, GC module 126, e.g., in conjunction with processor(s) 102, memory / memories 104, operating system 106, file system 120, etc., may perform GC of one or more candidate regions. In an example, GC module 126 may start with the first (e.g., highest ranked) candidate region from a subset of the plurality of regions, or may take a region with the highest garbage rate and outside of the residence time, and perform GC on that region, and then may move to the next region if the subset of the plurality of regions is determined to be a candidate region. As previously described, performing GC may include relocating valid data to another region (e.g., the last region added), which may include opening another region and then removing or deleting the region once only invalid data remains.
[0038] Figures 3A-3D is a flowchart of a specific example of a method 300 for determining candidate regions for GC in a file system according to aspects described herein. For example, the method 300 may be performed by the device 100 executing the file system 120.
[0039] refer to Figure 3A , in method 300, at 302, the flag run_gc_worker can be set to =TRUE. This can be a flag set by the file system 120 to activate the GC process. At 304, it can be determined (e.g., by the file system 120, the zone maintenance module 122, etc.) that Free_Percent>GC_START_LEVEL. For example, Free_Percent can be a measurement of the capacity of the block storage device 130, and GC_START_LEVEL can be a threshold for starting GC. If not, at 306, the GC is not triggered, and the thread executing the GC process can be awakened after sleeping for a certain period of time. If Free_Percent>GC_START_LEVEL, at 308, it can be determined (e.g., by the file system 120, the candidate determination module 124, etc.) that Free_Percent<=safe threshold and Free_Percent>=critical threshold. As previously described, for example, the safe threshold can be that if Free_Percent also meets at least the critical threshold, a safe GC process is performed (continue to Figure 3C and Figure 3D The critical threshold can be less than the safe threshold and can cause a critical GC process to be executed (continue to Figure 3B A).
[0040] refer to Figure 3B(e.g., for critical GC processes), at 310, GetFSSnapshot may be executed (e.g., by file system 120, candidate determination module 124, etc.) to obtain a file system snapshot (e.g., storing all information about extents in the current system state). This may include obtaining and storing (e.g., in memory) the garbage rate, sequence number, etc. of each extent in the file system at the time of the snapshot. At 312, the next closed extent may be considered (e.g., by file system 120, candidate determination module 124, etc.), which may be part of iterating through all extents from the snapshot. At 314, it may be determined (e.g., by file system 120, candidate determination module 124, etc.) whether scanning all extents is complete. If not, at 316, it may be determined (e.g., by file system 120, candidate determination module 124, etc.) whether the extent is below dwell_time. If so, the next closed extent may be considered at 312. If not (e.g., if the region is outside of dwell_time as described herein), then at 318, it may be determined (e.g., by the file system 120, the candidate determination module 124, etc.) whether the garbage rate of the region under consideration is greater than so_far_best_garbage_rate. If not, then at 312, the next closed region may be considered. If so, at 320, so_far_best_garbage_rate may be replaced (e.g., by the file system 120, the candidate determination module 124, etc.) with the garbage_rate of the region under consideration. If scanning all regions is complete at 314, then at 322, the region associated with so_far_best_garbage_rate may be set (e.g., by the file system 120, the candidate determination module 124, etc.) as a GC candidate. In this example, during a critical GC process, one region may be selected as a candidate for GC.
[0041] refer to Figure 3C (e.g., for a safe GC process), at 324, GetFSSnapshot can be executed (e.g., by the file system 120, the candidate determination module 124, etc.) to obtain a file system snapshot (e.g., storing all information about the zones in the current system state). This can include obtaining and storing (e.g., in memory) the garbage rate, sequence number, etc. of each zone in the file system at the snapshot time. At 326, the next closed zone can be considered (e.g., by the file system 120, the candidate determination module 124, etc.), which can be part of iterating through all zones from the snapshot. At 328, it can be determined (e.g., by the file system 120, the candidate determination module 124, etc.) whether scanning all zones is complete. If so, the method 300 can continue to Figure 3DIf not, then at 330, it may be determined (e.g., by the file system 120, the candidate determination module 124, etc.) whether the region is below dwell_time. If so, then at 326, the next closed region may be considered. If not (e.g., if the region is outside dwell_time as described herein), then at 332, it may be determined (e.g., by the file system 120, the candidate determination module 124, etc.) whether the garbage_rate is greater than the minimum garbage_rate of the potential candidates present in the top_candidate queue. If not (e.g., and the top_candidate queue is full), then at 326, the next closed region may be considered. If yes (e.g., and the top_candidate queue is full), then at 334, the potential GC candidate with the minimum garbage rate in the top_candidate queue may be replaced (e.g., by the file system 120, the candidate determination module 124, etc.) as the region under consideration. During and / or after this process, a top_candidate queue 336 is generated as an array arranged in descending order based on garbage_rate having a specific size.
[0042] refer to Figure 3D , if in Figure 3CIf all regions are scanned at 328 in the top_candidate queue, then at 338, the region at index 0 (e.g., the first region) in the top_candidate queue may have the max_garbage_rate. At 340, the next potential GC candidate may be obtained from the array (e.g., by the file system 120, the candidate determination module 124, etc.). At 342, it may be determined whether the file system 120, the candidate determination module 124, etc. has completed evaluating all candidates in the array (e.g., the top_candidate queue). If not, at 344, it may be determined (e.g., by the file system 120, the candidate determination module 124, etc.) whether the garbage_rate of the current candidate is at least the max_garbage_rate minus the garbage_rate_gap. If not, at 346, the file system 120, the candidate determination module 124, etc. may remove the candidate from the array and obtain the next potential GC candidate from the array at 340. If the garbage_rate of the current candidate is at least max_garbage_rate minus garbage_rate_gap, then the candidate may be maintained (e.g., by the file system 120, candidate determination module 124, etc.) for consideration in the array at 348. If the file system 120, candidate determination module 124, etc. finishes evaluating the candidates in the array at 342, then at 350, a minimum sequence number may be calculated from the extents in the array (e.g., by the file system 120, candidate determination module 124, etc.).
[0043] At 352, the remaining candidate areas may be traversed as follows. At 354, the file system 120, candidate determination module 124, etc. may obtain the next potential GC candidate from the array. At 356, a determination may be made as to whether the file system 120, candidate determination module 124, etc. has completed processing all candidate areas in the array. If not, at 358, a determination may be made (e.g., by the file system 120, candidate determination module 124, etc.) as to whether the sequence number of the candidate under consideration is less than or equal to the minimum sequence number in the array added to the sequence gap. If not, at 360, the candidate may be removed from the array (e.g., by the file system 120, candidate determination module 124, etc.). If so, at 362, the candidate may be retained in the array for consideration (e.g., by the file system 120, candidate determination module 124, etc.). If the file system 120, candidate determination module 124, etc. has completed processing all candidates in the array at 356, at 364, the best GC candidate in the array may be the candidate with the highest garbage rate among the remaining areas. In an example, at least the best GC candidate (and / or other GC candidates in the array) may be GCed. In an example, candidate regions from the array may be GCed in the order of the array.
[0044] Figure 4 An example of a device 400 is shown, which is similar to or similar to the device 100 ( Figure 1 ) are the same, including Figure 1 . In one embodiment, device 400 may include processor(s) 402, which may be similar to processor(s) 102, for performing processing functions associated with one or more of the components and functions described herein. Processor(s) 402 may include a single or multiple processor groups or multi-core processors. Furthermore, processor(s) 402 may be implemented as an integrated processing system and / or a distributed processing system.
[0045] The device 400 may also include memory / memories 404, which may be similar to the memory / memories 104, such as for storing local versions of applications executed by the processor(s) 402, the host application 110, the file system 120, related modules, instructions, parameters, etc. The memory / memories 404 may include a type of memory usable by a computer, such as RAM, ROM, tape, magnetic disk, optical disk, volatile memory, non-volatile memory, and any combination thereof.
[0046] Additionally, the device 400 may include a communication module 406 that provides for establishing and maintaining communications with one or more other devices, parties, entities, etc. utilizing hardware, software, and services as described herein. The communication module 406 may facilitate communications between modules on the device 400 and between the device 400 and external devices, such as devices located across a communication network and / or devices serially or locally connected to the device 400. For example, the communication module 406 may include one or more buses and may also include a transmit chain module and a receive chain module associated with a wireless or wired transmitter and receiver, respectively, operable to interface with external devices.
[0047] In addition, the device 400 may include a data store 408, which may be any suitable combination of hardware and / or software that provides mass storage of information, databases, and programs used in association with the embodiments described herein. For example, the data store 408 may be or may include a data repository for applications and / or related parameters (e.g., the host application 110, the file system 120, related modules, instructions, parameters, etc.) that are executed by the processor(s) 402 or are not currently being executed by the processor(s) 402. In one example, the data store 408 may include a block storage device 130. In addition, the data store 408 may be a data repository for the host application 110, the file system 120, related modules, instructions, parameters, etc., and / or one or more other modules of the device 400.
[0048] Device 400 may include a user interface module 410 that is operable to receive input from a user of device 400 and further operable to generate output for presentation to the user. User interface module 410 may include one or more input devices, including but not limited to a keyboard, a numeric keypad, a mouse, a touch-sensitive display, navigation keys, function keys, a microphone, a voice recognition component, a gesture recognition component, a depth sensor, a gaze tracking sensor, switches / buttons, any other mechanism capable of receiving input from a user, or any combination thereof. In addition, user interface module 410 may include one or more output devices, including but not limited to a display, a speaker, a tactile feedback mechanism, a printer, any other mechanism capable of presenting output to a user, or any combination thereof.
[0049] As an example, any combination of an element or any part of an element or an element can be implemented with a "processing system" comprising one or more processors. The example of a processor includes a microprocessor, a microcontroller, a digital signal processor (DSP), a field programmable gate array (FPGA), a programmable logic device (PLD), a state machine, a gated logic, a discrete hardware circuit and other suitable hardware configured to perform the various functions described throughout this disclosure. One or more processors in a processing system can execute software. Software should be broadly interpreted as referring to instructions, instruction sets, codes, code segments, program codes, programs, subroutines, software modules, applications, software applications, software packages, routines, subroutines, objects, executable files, execution threads, processes, functions, etc., whether referred to as software, firmware, middleware, microcode, hardware description language or other.
[0050] Therefore, in one or more embodiments, one or more of the functions described can be implemented with hardware, software, firmware, or any combination thereof. If implemented with software, the function can be stored on a computer-readable medium or encoded as one or more instructions or codes on a computer-readable medium. Computer-readable media include computer storage media. The storage medium can be any available medium that can be accessed by a computer. As an example and not limitation, such computer-readable media can include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, disk storage or other magnetic storage devices, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and can be accessed by a computer. As used herein, disks and optical disks include compact disks (CDs), laser disks, optical disks, digital versatile disks (DVDs), and floppy disks, wherein disks typically reproduce data magnetically, while optical disks reproduce data optically with lasers. The above combinations should also be included within the scope of computer-readable media.
[0051] The foregoing description is provided to enable any person skilled in the art to practice the various embodiments described herein. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other embodiments. Therefore, the claims are not intended to be limited to the embodiments shown herein, but are to be given the full scope consistent with the claim language, wherein reference to a singular element is not intended to mean "one and only one" unless specifically stated otherwise, but rather "one or more". Unless otherwise specifically stated, the term "some" refers to one or more. All structural and functional equivalents of the elements of the various embodiments described herein, which are known or will later become known to those of ordinary skill in the art, are intended to be covered by the claims. In addition, nothing disclosed herein is intended to be dedicated to the public, regardless of whether such disclosure is explicitly stated in the claims. Claim elements are not to be interpreted as means plus function unless an element is explicitly stated as "means for..."
Claims
1. A method for performing garbage collection implemented on a computer, for a file system having a plurality of data areas, wherein the method evaluates the hotness or coldness of the data by assigning sequence numbers to the plurality of areas, comprising: For each area in the file system, calculating a garbage rate of each area, wherein the garbage rate is related to the amount of invalid data in the area; determining one or more candidate regions among the plurality of regions for prioritizing garbage collection for the candidate regions based on the garbage rate and a sequence number assigned to each region, wherein the sequence number is assigned to each region when the region is created to store data; as well as Garbage collection is performed on the one or more candidate areas in the file system.
2. The method of claim 1 , wherein determining the one or more candidate regions comprises: Determining that the sequence number of the one or more candidate regions exceeds a residence time range of a highest sequence number among the plurality of regions, wherein the residence time represents a range of sequence numbers from a highest sequence number corresponding to a newly allocated region to a sequence number to be avoided during garbage collection.
3. The method of claim 1 , wherein the one or more candidate regions include a plurality of candidate regions, and wherein determining the plurality of candidate regions comprises: determining a subset of the plurality of candidate areas having a highest garbage rate among the one or more candidate areas; as well as removing at least one candidate region from a subset of one or more candidate regions having a highest junk rate, the junk rate of the candidate region being one junk rate gap lower than a highest junk rate of a highest candidate region in the subset; The garbage rate gap is configured by the file system and represents a garbage rate range starting from the highest garbage rate to be considered in garbage collection. 4 . The method of claim 3 , wherein performing the determining of the plurality of candidate areas is based at least in part on determining that a capacity of the file system reaches a threshold.
5. The method of claim 1 , wherein the one or more candidate regions include a plurality of candidate areas, and wherein determining the plurality of candidate regions comprises: determining a subset comprising a plurality of regions having a highest garbage rate among the one or more candidate regions; as well as At least one candidate region having a higher sequence number than other candidate regions is removed from the subset to prioritize regions with colder data for garbage collection.
6. The method of claim 1 , wherein the one or more candidate regions include a plurality of candidate regions, and wherein determining the plurality of candidate regions includes: determining a subset comprising a plurality of regions with the highest garbage rates among the one or more candidate regions; as well as Removing at least one candidate area from a subset of the plurality of candidate areas among the one or more candidate areas, the at least one candidate area having a sequence number greater than a sequence number of a candidate area with a lowest sequence number in the subset by a sequence number gap, wherein the sequence number gap is configured by the file system and represents a range of sequence numbers starting from the lowest sequence number to be considered for garbage collection.
7. The method of claim 1 , wherein determining the one or more candidate regions comprises determining the one or more candidate regions having the highest garbage rate and having sequence numbers that are outside a residence time from a highest sequence number among the plurality of regions, wherein the residence time represents a range of sequence numbers starting from the highest sequence number to be avoided during garbage collection, the highest sequence number corresponding to the most recently allocated region.
8. The method of claim 1, wherein determining the one or more candidate areas comprises determining that the one or more candidate areas are opened by a host application such that areas opened by the file system for overwriting valid data are not considered for garbage collection.
9. A device for performing garbage collection in a file system having multiple data areas, the device comprising one or more processors and one or more non-volatile memories having instructions thereon, wherein the instructions, when executed by the one or more processors, cause the one or more processors to perform the method described in any one of claims 1 to 9.
10. One or more non-transitory computer-readable storage media storing instructions that, when executed by one or more processors, cause the one or more processors to perform the method of any one of claims 1 to 8.