Drive-Based Storage Scrubbing for Low-Overhead Data Integrity

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional data storage systems face substantial overhead during scrubbing operations, especially in distributed environments using disaggregated storage, as they require reading and validating large amounts of data over networks, which is inefficient and resource-intensive.

Innovation Solution

Shifting the scrubbing process from host-based to drive-based validation, where the scrubber is located locally with the storage media, allowing only error data to be transferred over the network for verification and correction, thereby reducing the need for extensive data transfer and processing.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If host-based scrubbing is used to validate stored data, then data integrity can be ensured, but substantial overhead is generated due to reading and validating large amounts of data over networks

Engineering Contradiction:
Improvedata integrityVSAvoidscrubbing overhead
Core Design Contradiction:
ReliabilityVSLoss of energy

Solution Approach 1:

The patent extracts the scrubbing function from the host system and places it directly on the storage device. The storage device now performs self-validation using its own scrubbing circuitry, eliminating the need for the host to read and validate large amounts of data over the network. Only validation results are transmitted back to the host, dramatically reducing network traffic and host processing overhead while maintaining data integrity verification.

Inventive Principle:
Principle #2Taking out (Extraction)

2Adaptability or versatility

If disaggregated storage is used to distribute data across networks, then storage capacity and accessibility are improved, but the amount of data that must be accessed over networks during scrubbing increases substantially

Engineering Contradiction:
Improvestorage distributionVSAvoiddata transfer volume
Core Design Contradiction:
Adaptability or versatilityVSQuantity of substance

Solution Approach 1:

The storage device performs self-service by implementing onboard scrubbing capability. The storage device independently validates its own stored data using integrated scrubbing circuitry and logic, without requiring the host to retrieve and validate the data. This self-validation approach maintains the benefits of disaggregated storage distribution while eliminating the substantial network data transfer that would otherwise be required during scrubbing operations.

Inventive Principle:
Principle #25Self-service

3Reliability

If all stored data is read and validated by the host during scrubbing, then comprehensive data validation is achieved, but resource utilization becomes inefficient

Engineering Contradiction:
Improvedata validation completenessVSAvoidscrubbing efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The validation function is extracted from the host and implemented on the storage device itself. The storage device's scrubbing circuitry performs comprehensive validation of all stored data locally, ensuring complete data validation coverage. Only the results of this validation are sent to the host, transforming an inefficient host-centric validation process into an efficient storage-device-centric process that maintains completeness while dramatically improving productivity.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent shifts the validation operation from the host dimension to the storage device dimension. Instead of the host reading and validating data across the network, the validation occurs in-place at the storage device level. This dimensional shift in where validation occurs enables comprehensive data validation without the network overhead, resolving the contradiction between validation completeness and scrubbing efficiency.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS10394634B2Drive-based storage scrubbing
Publication Date: 2019.08.27 INTEL CORP
  • US10394634B2 patent drawing
  • US10394634B2 patent drawing
  • US10394634B2 patent drawing

AI summary

Apparatuses, systems and methods are disclosed herein that generally relate to distributed storage, such as for big data, distributed databases, large datasets, artificial intelligence, genomics, or any other data processing environment using that host large data sets or utilize big data hosts using local storage or storage remotely located over a network. More particularly since large scale data requires many storage devices, scrubbing storage for reliability and accuracy requires communication bandwidth and processor resources. Discussed are various ways to use known storage structure, such as LBA, to offload scrubbing overhead to storage by having storage engage in autonomous self-validation. Storage may scrub itself and identify stored data failing data integrity validation, or identify unreadable storage locations, and report errors to a distributed storage system that may reverse-lookup the affected storage location to identify, for example, a data block at that location needing correction.