File-System Zone Garbage Collection With Age-Aware Selection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current garbage collection algorithms in file systems fail to consider data hotness and coldness, leading to increased write amplification and lack of segregation between hot and cold data, as they prioritize zones based solely on garbage rate without considering recent data access patterns.
Innovation Solution
Implement a method for garbage collection that computes a garbage rate and sequence number for each zone, using dwell time and threshold-based criteria to select candidate zones for garbage collection, ensuring hot data is not collected and maintaining segregation between hot and cold data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If current GC algorithms select zones based solely on maximum garbage rate, then the maximum space can be freed by GCing zones with the highest amount of invalid data, but write amplification increases because zones with hot data are selected for GC and subsequently rewritten by the host
Solution Approach 1:
The patent changes the selection parameter from only garbage rate to a composite criterion including garbage rate and sequence number (representing data age). This parameter change allows the system to avoid selecting zones with hot data (recent sequence numbers) while still identifying zones with high garbage rates, thereby reducing write amplification while maintaining space reclamation effectiveness
Solution Approach 2:
The patent introduces a dwell time mechanism that prevents zones from being selected for GC immediately after they are written to. By requiring zones to remain in the pool for a minimum dwell time, the system preliminarily filters out zones containing hot data before GC selection, preventing subsequent rewrites and reducing write amplification
2Productivity
If current GC algorithms prioritize zones with maximum invalid data, then GC efficiency is improved, but segregation between hot and cold data is lost causing increased write amplification
Solution Approach 1:
The patent applies local quality by treating zones with different sequence numbers (representing different ages) differently in the GC selection process. Zones with recent sequence numbers are excluded or deprioritized, while zones with older sequence numbers are preferentially selected. This local differentiation maintains data segregation between hot and cold data while preserving GC efficiency
3Device complexity
If GC algorithms select zones without considering sequence numbers, then simple implementation is maintained, but hot data is inadvertently collected leading to increased write amplification
Solution Approach 1:
The patent performs preliminary action by pre-calculating and storing sequence numbers for each zone when data is written. This advance preparation allows the GC algorithm to efficiently query sequence numbers during zone selection without adding significant computational overhead, maintaining simplicity while enabling hot data identification and exclusion
Data Source
Figure 1
Figure 2
Figure 3A
AI summary
Described are examples for performing garbage collection in a file system having multiple zones of data. A garbage rate associated with an amount of invalid data in the zone can be computed for each zone of the multiple zones in the file system. One or more candidate zones, of the multiple zones, can be determined for garbage collection based on the garbage rate and a sequence number assigned to the zone. Garbage collection of the one or more candidate zones in the file system can be performed.