Ghost Fingerprint Deduplication for NAND Flash Storage
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Data deduplication in NAND flash memory systems faces challenges in managing fingerprints efficiently, leading to increased wear and reduced performance due to resource-intensive background scanning and limited dynamic memory capacity, which impacts overall I/O performance.
Innovation Solution
A controller in the data storage system generates fingerprints for data blocks and maintains state information to selectively perform deduplication, caching frequently accessed fingerprints and lazily managing ghost fingerprints to reduce metadata overhead and improve deduplication efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of time
If background data deduplication is employed, then store latency is reduced, but storage capacity requirement and wear on storage media increase
Solution Approach 1:
The patent segments the deduplication process into two distinct modes: background deduplication for reducing store latency, and inline deduplication for reducing storage capacity requirements. The system dynamically selects which mode to use based on current operational conditions, allowing it to benefit from both approaches without suffering their respective drawbacks simultaneously.
Solution Approach 2:
The patent implements dynamic switching between background and inline deduplication modes based on real-time system state. The controller monitors storage capacity availability and I/O workload characteristics, adapting the deduplication strategy accordingly. This dynamic approach allows the system to optimize for store latency when capacity is abundant, and optimize for capacity utilization when storage is constrained.
2Quantity of substance
If in-line data deduplication is employed, then storage capacity and wear are reduced, but store latency and write bandwidth decrease
Solution Approach 1:
The system dynamically switches between inline and background deduplication modes based on real-time conditions. When storage capacity is constrained or I/O performance is prioritized, the system uses background deduplication. When capacity is abundant and deduplication ratio is low, the system employs inline deduplication to minimize storage usage.
Solution Approach 2:
The patent changes the operational parameters of the deduplication process based on system state. It monitors metrics such as storage capacity availability, deduplication ratio, and I/O workload characteristics, then adjusts the deduplication mode accordingly. This parameter-based control allows optimal performance under varying conditions.
3Productivity
If a large amount of dynamic memory is used for fingerprint storage, then I/O performance is improved, but memory cost and complexity increase
Solution Approach 1:
The patent introduces a fingerprint cache as an intermediary layer between the fingerprint index (stored in non-volatile memory) and the I/O processing path. Frequently accessed fingerprints are cached in dynamic memory, providing fast access for hot data while avoiding the need to store all fingerprints in expensive high-speed memory. This cache-mediated approach delivers high I/O performance for common operations while keeping overall memory requirements manageable.
4Quantity of substance
If background scanning is performed for deduplication, then deduplication ratio is improved, but processing overhead and wear on storage media increase
Solution Approach 1:
The patent implements periodic background scanning instead of continuous scanning. The background deduplication process is triggered at scheduled intervals or based on accumulated data volume, rather than continuously processing every I/O operation. This periodic approach achieves good deduplication ratios over time while significantly reducing the sustained processing overhead and wear compared to continuous background scanning.
Data Source
AI summary
A controller of a data storage system generates fingerprints of data blocks written to the data storage system. The controller maintains, in a data structure, respective state information for each of a plurality of data blocks. The state information for each data block can be independently set to indicate any of a plurality of states, including at least one deduplication state and at least one non-deduplication state. At allocation of a data block, the controller initializes the state information for the data block to a non-deduplication state and, thereafter, in response to detection of a write of duplicate of the data block to the data storage system, transitions the state information for the data block to a deduplication state. The controller selectively performs data deduplication for data blocks written to the data storage system based on the state information in the data structure and by reference to the fingerprints.


