Efficiency Sets for Sampling Data Characteristics in Distributed Storage

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Distributed storage systems using a content-driven approach face challenges in monitoring and managing data characteristics due to limited visibility and complexity in obtaining performance metrics, leading to inefficiencies in system management and operation.

Innovation Solution

A statistical sampling method is employed to create an efficiency set of block IDs, allowing for the estimation of data characteristics such as compression state, entropy, and data type by analyzing metadata and data within a subset of block IDs, which can be extrapolated to the entire data set.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of substance

If a content-driven approach is used to store data blocks with the same content having the same block ID, then storage efficiency is improved through deduplication, but performance monitoring and system management functions become more complex

Engineering Contradiction:
Improvestorage efficiencyVSAvoidsystem management complexity
Core Design Contradiction:
Loss of substanceVSDevice complexity

Solution Approach 1:

The patent introduces an intermediary process that samples block IDs from the content-driven storage system and retrieves their metadata characteristics. This intermediary layer translates the complex internal storage structure into simplified performance metrics that can be monitored without directly accessing the complex block ID mapping system.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent replaces direct mechanical access to storage metadata with a statistical sampling approach. Instead of systematically accessing and analyzing all block IDs (which would be complex), the system uses random sampling to estimate performance characteristics, substituting a simpler statistical mechanism for the complex direct access method.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If all block ID metadata is analyzed to obtain performance metrics, then measurement precision is improved, but loss of time and computational resources increases

Engineering Contradiction:
Improveperformance metric accuracyVSAvoidtime to obtain metrics
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent applies partial action by sampling only a subset of block IDs rather than analyzing all blocks. The sampling process selects representative blocks randomly, obtaining sufficient performance metric accuracy without the time cost of comprehensive analysis of every block in the storage system.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The system performs preliminary sampling to identify representative blocks before conducting detailed metadata analysis. This preliminary selection step allows the system to focus computational resources on a manageable subset of blocks that adequately represent the overall storage characteristics.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12386507B2Creation and use of an efficiency set to estimate an amount of data stored in a data set of a storage system having one or more characteristics
Publication Date: 2025.08.12 NETAPP INC
  • US12386507B2 patent drawing
  • US12386507B2 patent drawing
  • US12386507B2 patent drawing

AI summary

Systems and methods for sampling a set of block IDs to facilitate estimating an amount of data stored in a data set of a storage system having one or more characteristics are provided. According to an example, metadata (e.g., block headers and block IDs) may be maintained regarding multiple data blocks of the data set. When one or more metrics relating to the data set are desired, an efficiency set, representing a subset of the block IDs of the data set, may be created to facilitate efficient calculation of the metrics by sampling the block IDs of the data set. Finally, the metrics may be estimated based on the efficiency set by analyzing one or more of the metadata (e.g., block headers) and the data contained in the data blocks corresponding to the subset of the block IDs and extrapolating the metrics for the entirety of the data set.