Dynamic RAID Level Assignment for De-duplicated Storage Integrity
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
De-duplicated storage environments face challenges in maintaining data integrity due to the increased vulnerability of data corruption and disk failures, as higher RAID levels required for protection contradict the purpose of de-duplication by necessitating more storage.
Innovation Solution
Implementing dynamic software-defined RAID levels that adjust based on de-duplication information, where highly de-duplicated data is placed on higher RAID levels and less de-duplicated data on lower levels, along with dynamic data placement between performance pools, to balance protection and storage efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If higher RAID levels are used to protect de-duplicated data, then data integrity is improved, but storage efficiency deteriorates due to increased storage requirements
Solution Approach 1:
The patent applies local quality by assigning different RAID protection levels to different data chunks based on their individual de-duplication counts. Highly de-duplicated chunks (with higher importance) are placed on higher RAID levels for enhanced protection, while less de-duplicated chunks are placed on lower RAID levels. This localized differentiation resolves the contradiction by providing targeted protection only where needed, rather than uniformly applying high RAID levels to all data.
Solution Approach 2:
The patent changes the RAID level parameter dynamically based on the de-duplication count parameter. By monitoring how many times each data chunk has been de-duplicated and adjusting the RAID protection level accordingly, the system adapts the storage protection to match the actual importance and redundancy of each chunk, thereby optimizing both data integrity and storage efficiency.
2Reliability
If uniform high RAID levels are applied to all de-duplicated data, then data integrity is improved, but storage capacity is reduced
Solution Approach 1:
Instead of applying uniform high RAID levels across all storage, the patent implements local quality by differentiating protection levels based on individual chunk characteristics. Only chunks with high de-duplication counts receive enhanced RAID protection, while other chunks use standard or reduced protection, thereby preserving overall storage capacity while maintaining data integrity for critical chunks.
Solution Approach 2:
The patent applies partial action by providing enhanced RAID protection to only those data chunks that require it (those with high de-duplication counts), rather than applying excessive protection uniformly to all data. This selective approach ensures adequate data protection where needed while avoiding the storage capacity consumption that would result from universal high-level protection.
3Quantity of substance
If de-duplication is increased to improve storage efficiency, then storage capacity is optimized, but data vulnerability to corruption increases
Solution Approach 1:
The patent changes the RAID protection parameter based on the de-duplication parameter. As de-duplication increases for a given chunk (indicating higher storage efficiency gains), the system responds by increasing the RAID protection level for that chunk. This dynamic parameter adjustment compensates for the increased corruption vulnerability introduced by higher de-duplication while maintaining overall storage efficiency.
Solution Approach 2:
The patent applies preliminary anti-action by proactively increasing RAID protection for data chunks that have been highly de-duplicated. By detecting high de-duplication counts and preemptively assigning higher RAID levels, the system counteracts the increased vulnerability to data corruption that arises from aggressive de-duplication, thereby protecting against potential harm before it can manifest.
Data Source
AI summary
A mechanism is provided in a data processing system for securing data integrity in de-duplicated storage environments in combination with software defined native redundant array of independent disks (RAID). The mechanism receives a data portion to write to storage, divides the data portion into a plurality of chunks, and identifies a given chunk within the plurality of chunks for de-duplication. The mechanism increment a de-duplication counter for the given chunk and determines a RAID level for the given chunk based on a value of the de-duplication counter. The mechanism stores the given chunk based on the determined RAID level.


