Drive-Based Storage Scrubbing for Low-Overhead Data Integrity
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional data storage systems face substantial overhead during scrubbing operations, especially in distributed environments using disaggregated storage, as they require reading and validating large amounts of data over networks, which is inefficient and resource-intensive.
Innovation Solution
Shifting the scrubbing process from host-based to drive-based validation, where the scrubber is located locally with the storage media, allowing only error data to be transferred over the network for verification and correction, thereby reducing the need for extensive data transfer and processing.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If host-based scrubbing is used to validate stored data, then data integrity can be ensured, but substantial overhead is generated due to reading and validating large amounts of data over networks
Solution Approach 1:
The patent extracts the scrubbing function from the host system and places it directly on the storage device. The storage device now performs self-validation using its own scrubbing circuitry, eliminating the need for the host to read and validate large amounts of data over the network. Only validation results are transmitted back to the host, dramatically reducing network traffic and host processing overhead while maintaining data integrity verification.
2Adaptability or versatility
If disaggregated storage is used to distribute data across networks, then storage capacity and accessibility are improved, but the amount of data that must be accessed over networks during scrubbing increases substantially
Solution Approach 1:
The storage device performs self-service by implementing onboard scrubbing capability. The storage device independently validates its own stored data using integrated scrubbing circuitry and logic, without requiring the host to retrieve and validate the data. This self-validation approach maintains the benefits of disaggregated storage distribution while eliminating the substantial network data transfer that would otherwise be required during scrubbing operations.
3Reliability
If all stored data is read and validated by the host during scrubbing, then comprehensive data validation is achieved, but resource utilization becomes inefficient
Solution Approach 1:
The validation function is extracted from the host and implemented on the storage device itself. The storage device's scrubbing circuitry performs comprehensive validation of all stored data locally, ensuring complete data validation coverage. Only the results of this validation are sent to the host, transforming an inefficient host-centric validation process into an efficient storage-device-centric process that maintains completeness while dramatically improving productivity.
Solution Approach 2:
The patent shifts the validation operation from the host dimension to the storage device dimension. Instead of the host reading and validating data across the network, the validation occurs in-place at the storage device level. This dimensional shift in where validation occurs enables comprehensive data validation without the network overhead, resolving the contradiction between validation completeness and scrubbing efficiency.
Data Source
AI summary
Apparatuses, systems and methods are disclosed herein that generally relate to distributed storage, such as for big data, distributed databases, large datasets, artificial intelligence, genomics, or any other data processing environment using that host large data sets or utilize big data hosts using local storage or storage remotely located over a network. More particularly since large scale data requires many storage devices, scrubbing storage for reliability and accuracy requires communication bandwidth and processor resources. Discussed are various ways to use known storage structure, such as LBA, to offload scrubbing overhead to storage by having storage engage in autonomous self-validation. Storage may scrub itself and identify stored data failing data integrity validation, or identify unreadable storage locations, and report errors to a distributed storage system that may reverse-lookup the affected storage location to identify, for example, a data block at that location needing correction.


