Log-Structured Merge Tree Storage Using Dual Filters for SSD I/O

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing key-value storage engines, such as RocksDB, suffer from high context-switch overhead due to synchronous file I/O operations, which degrades storage performance despite low SSD latency, and require costly multi-channel mechanisms to support multiple SSDs with limited CPU resources.

Innovation Solution

A log-structured merge tree-based data storage system using a bloom filter and tag filter to efficiently locate target keys within SST files, allowing for standardized storage units and reduced memory usage, enabling parallel read and write operations across multiple SSDs with limited CPU resources.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If synchronous file I/O operations are used to read key-value pairs, then data can be read from storage devices, but the calling thread must wait for file I/O to complete, triggering context-switch overhead that degrades storage performance

Engineering Contradiction:
Improvestorage performanceVSAvoidcontext-switch overhead
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent pre-loads filter data (bloom filter and tag filter) into memory before actual query operations. This preliminary action allows the system to quickly determine whether target keys exist in SST files without performing full file I/O operations, thereby reducing context-switch overhead and improving storage performance by avoiding unnecessary synchronous waits.

Inventive Principle:
Principle #10Preliminary action

2Productivity

If multiple CPUs or CPU cores are used to perform read and write operations on multiple SSD storage devices, then read and write delays are reduced, but the cost increases and CPU resources are consumed

Engineering Contradiction:
Improveread and write throughputVSAvoidmulti-channel mechanism
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent introduces filter data (bloom filter and tag filter) as intermediaries between the query system and SST files. These filters act as a preliminary screening layer that quickly identifies whether target keys exist in specific SST files, reducing the need for full file scans and enabling more efficient single-threaded or limited-threaded access patterns, thereby reducing the need for complex multi-CPU mechanisms.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Ease of operation

If traditional indexing methods are used to locate target keys in SST files, then key lookup can be performed, but memory usage increases

Engineering Contradiction:
Improvekey lookup efficiencyVSAvoidmemory usage
Core Design Contradiction:
Ease of operationVSQuantity of substance

Solution Approach 1:

The patent segments the indexing structure into two parts: a bloom filter that provides coarse-grained key existence checking, and a tag filter that provides fine-grained location information. This segmentation allows the system to use minimal memory for filtering while maintaining efficient key lookup, as the bloom filter quickly eliminates non-matching keys and the tag filter only needs to store compact location identifiers rather than full index entries.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS20240220460A1Data storage device and storage control method based on log-structured merge tree
Publication Date: 2024.07.04 HONEYCOMBDATA INC
  • US20240220460A1 patent drawing
  • US20240220460A1 patent drawing
  • US20240220460A1 patent drawing

AI summary

The present invention relates to a data storage device and a storage control method based on a log-structured merge tree. The log-structured merge tree comprises a plurality of SST files stored on at least one storage device. The storage control method uses a standardized storage unit to store key-value pairs and uses two different filters to locate the storage units storing the target key-value pair in the SST files, thereby saving memory usage and improving file IO efficiency.