Bucket-Partitioned Hash Table Cache Management

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current caching techniques in data storage systems are inefficient due to the large size of cache metadata structures, which occupy significant space and reduce the capacity for storing actual user data, as the number of cache entries increases with the size of the cache, leading to suboptimal performance in managing and retrieving data.

Innovation Solution

The proposed solution involves partitioning the cache metadata into buckets, where each bucket has a dedicated section of the cache for exclusive use, using a hash table to determine bucket allocation based on a hash value generated from a data block's key, and storing metadata and data blocks in specific cache locations within these buckets, thereby reducing the size of pointer fields and optimizing cache usage.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If the cache size is increased to store more user data, then the cache capacity is improved, but the cache metadata size increases proportionally, occupying more space and reducing the effective storage capacity

Engineering Contradiction:
Improvecache capacityVSAvoidmetadata size
Core Design Contradiction:
Quantity of substanceVSDevice complexity

Solution Approach 1:

The hash table is divided into multiple buckets, where each bucket manages a specific section of the cache. This segmentation allows the metadata to be distributed across multiple smaller structures rather than one large structure, reducing the overhead per bucket and improving cache utilization efficiency.

Inventive Principle:
Principle #1Segmentation

2Quantity of substance

If the number of cache entries is increased to improve data storage capacity, then the cache capacity is improved, but the metadata structure becomes larger and more complex, leading to suboptimal performance in managing and retrieving data

Engineering Contradiction:
Improvenumber of cache entriesVSAvoiddata retrieval performance
Core Design Contradiction:
Quantity of substanceVSProductivity

Solution Approach 1:

By dividing the hash table into multiple buckets, each bucket can be managed independently with its own section of the cache. This segmentation improves data retrieval performance by allowing parallel access to different buckets and reducing the time required to search and manage metadata, even as the total number of cache entries increases.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Each bucket is assigned a dedicated section of the cache with specific characteristics optimized for that bucket's data. This local optimization allows for more efficient data management and retrieval within each bucket, improving overall system performance as the cache scales.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS11210231B2Cache management using a bucket-partitioned hash table
Publication Date: 2021.12.28 EMC IP HLDG CO LLC
  • US11210231B2 patent drawing
  • US11210231B2 patent drawing
  • US11210231B2 patent drawing

AI summary

Techniques for performing cache management includes partitioning entries of a hash table into buckets, wherein each of the buckets includes a portion of the entries of the hash table, configuring a cache, wherein the configuring includes allocating a section of the cache for exclusive use by each bucket, and performing first processing that stores a data block in the cache. The first processing includes determining a hash value for a data block, selecting, in accordance with the hash value, a first bucket of the plurality of buckets, wherein a first section of the cache is used exclusively for storing cached data blocks of the first bucket, storing metadata used in connection with caching the data block in a first entry of the first bucket, and storing the data block in a first cache location of the first section of the cache.