Granular Disk Unit Rebuilding for Partial Storage Failures
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
As storage systems with increasing disk capacities face partial disk failures, existing data rebuilding methods are inefficient, requiring extensive computing and network resources, and often result in unnecessary waste and reduced system responsiveness due to the need to rebuild all data, even when only a portion of the disk is faulty.
Innovation Solution
Implementing a system management-level data rebuilding process that identifies and rebuilds only the failed data block, allowing other healthy disk units to remain accessible and continue using the remaining storage space without going offline, thereby reducing resource utilization and maintaining system performance.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional data rebuilding method is used to rebuild all data in a failed disk, then data reliability is improved, but computing resources and time are excessively consumed
Solution Approach 1:
The patent segments the disk into multiple disk units (DU) and implements granular failure detection at the DU level rather than treating the entire disk as a single unit. This allows the system to identify and rebuild only the specific failed DU while leaving other healthy DUs accessible, thereby improving rebuilding efficiency without compromising data reliability.
Solution Approach 2:
The patent extracts and isolates the failed disk unit from the rest of the disk system. By identifying the specific failed DU through health status information and mapping mechanisms, the system separates the failure impact to only the affected DU, allowing healthy portions of the disk to remain operational and reducing the scope of rebuilding operations.
2Reliability
If traditional data rebuilding method is used to rebuild all data in a failed disk, then data protection is improved, but network resources are excessively occupied
Solution Approach 1:
The patent segments the rebuilding operation to only include the failed disk unit rather than the entire disk. This segmentation reduces the volume of data that needs to be transmitted over the network during rebuilding operations, thereby reducing network resource consumption while maintaining adequate data protection for the affected portion.
Solution Approach 2:
The patent applies partial action by performing rebuilding operations only on the failed disk unit rather than the entire disk. This partial rebuilding approach consumes fewer network resources compared to full disk rebuilding, while still providing sufficient data protection for the failed portion through targeted recovery operations.
3Reliability
If the entire disk is taken offline for rebuilding, then data safety is improved, but storage capacity and system responsiveness are reduced
Solution Approach 1:
The patent segments the disk into multiple independent disk units and implements failure isolation at the DU level. This allows the system to take only the failed DU offline for rebuilding while keeping other healthy DUs online and accessible, thereby maintaining storage capacity and system responsiveness while ensuring data safety for the affected portion.
Solution Approach 2:
The patent applies partial action by offline-ing only the failed disk unit rather than the entire disk. This partial offline approach preserves accessibility to healthy storage portions, maintaining system responsiveness and usable storage capacity while still ensuring data safety through targeted rebuilding operations on the failed unit.
Data Source
AI summary
Techniques provide for rebuilding data. Such techniques involve: obtaining health status information related to a first disk of a storage system, the first disk being divided into a plurality of disk units, and the health status information indicating a failure of a first disk unit of the plurality of disk units; determining a data block stored in the first disk unit based on a mapping between data blocks for the storage system and storage locations; and rebuilding the data block into a second disk of the storage system when maintaining accessibility of other data blocks in other disk units of the first disk than the first disk unit. Accordingly, it is possible to improve the data rebuilding efficiency when a disk fails partly and to continue utilizing the storage space portion in the disk that is not failed, without making the disk be offline temporarily.


