Storage Control for De-Duplication With Regenerated Error Codes
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data de-duplication techniques in storage systems face challenges in maintaining data reliability when error detection codes are separated and reattached, leading to decreased efficiency and reliability, especially when data is compressed or changed.
Innovation Solution
A storage control apparatus that separates error detection codes from target data, eliminates duplicates, generates new error detection codes for de-duplicated data, and writes these codes back into the memory device, ensuring data integrity and capacity efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If error detection codes are separated from data to enable de-duplication, then capacity utilization is improved, but data reliability deteriorates
Solution Approach 1:
The patent segments the data processing into two distinct phases: first separating error detection codes from target data to enable de-duplication, then regenerating error detection codes from the de-duplicated data. This segmentation allows de-duplication to work on pure data while reliability is maintained through code regeneration, resolving the contradiction between capacity utilization and data reliability.
Solution Approach 2:
The patent performs preliminary separation of error detection codes before de-duplication, then performs the action of regenerating new error detection codes after de-duplication. This preliminary and subsequent action sequence ensures that de-duplication operates on clean data without codes, maximizing capacity utilization while the subsequent code regeneration restores data reliability.
2Reliability
If de-duplication is performed on data with error detection codes, then data reliability is maintained, but capacity utilization deteriorates
Solution Approach 1:
The patent extracts error detection codes from the target data before performing de-duplication. By taking out the codes that cause de-duplication inefficiency, the system can perform de-duplication on pure data, achieving high capacity utilization. The extracted codes are then used to verify original data while new codes are generated for de-duplicated data, maintaining reliability.
3Reliability
If error detection codes are regenerated after de-duplication, then data reliability is improved, but processing complexity increases
Solution Approach 1:
The system performs self-service by automatically regenerating error detection codes from the de-duplicated data using the same de-duplication apparatus. This self-service mechanism improves data reliability without requiring external intervention or significantly increasing processing complexity, as the code generation uses existing de-duplication infrastructure.
Data Source
AI summary
A storage control apparatus obtains first-code attached data, each having target data to be written and first code information, which includes an error detection code based on the target data and information about a first write destination, attached to the target data. The storage control apparatus then obtains the target data by excluding the first code information from the first-code attached data eliminates duplication of the target data, generates second code information which includes an error detection code for the target data remaining and information about a second write destination, and writes second-code attached data including the second code information into a memory device.


