Log-Structured Merge Tree Storage Using Dual Filters for SSD I/O
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing key-value storage engines, such as RocksDB, suffer from high context-switch overhead due to synchronous file I/O operations, which degrades storage performance despite low SSD latency, and require costly multi-channel mechanisms to support multiple SSDs with limited CPU resources.
Innovation Solution
A log-structured merge tree-based data storage system using a bloom filter and tag filter to efficiently locate target keys within SST files, allowing for standardized storage units and reduced memory usage, enabling parallel read and write operations across multiple SSDs with limited CPU resources.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If synchronous file I/O operations are used to read key-value pairs, then data can be read from storage devices, but the calling thread must wait for file I/O to complete, triggering context-switch overhead that degrades storage performance
Solution Approach 1:
The patent pre-loads filter data (bloom filter and tag filter) into memory before actual query operations. This preliminary action allows the system to quickly determine whether target keys exist in SST files without performing full file I/O operations, thereby reducing context-switch overhead and improving storage performance by avoiding unnecessary synchronous waits.
2Productivity
If multiple CPUs or CPU cores are used to perform read and write operations on multiple SSD storage devices, then read and write delays are reduced, but the cost increases and CPU resources are consumed
Solution Approach 1:
The patent introduces filter data (bloom filter and tag filter) as intermediaries between the query system and SST files. These filters act as a preliminary screening layer that quickly identifies whether target keys exist in specific SST files, reducing the need for full file scans and enabling more efficient single-threaded or limited-threaded access patterns, thereby reducing the need for complex multi-CPU mechanisms.
3Ease of operation
If traditional indexing methods are used to locate target keys in SST files, then key lookup can be performed, but memory usage increases
Solution Approach 1:
The patent segments the indexing structure into two parts: a bloom filter that provides coarse-grained key existence checking, and a tag filter that provides fine-grained location information. This segmentation allows the system to use minimal memory for filtering while maintaining efficient key lookup, as the bloom filter quickly eliminates non-matching keys and the tag filter only needs to store compact location identifiers rather than full index entries.
Data Source
AI summary
The present invention relates to a data storage device and a storage control method based on a log-structured merge tree. The log-structured merge tree comprises a plurality of SST files stored on at least one storage device. The storage control method uses a standardized storage unit to store key-value pairs and uses two different filters to locate the storage units storing the target key-value pair in the SST files, thereby saving memory usage and improving file IO efficiency.


