Storage Control for De-Duplication With Regenerated Error Codes

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing data de-duplication techniques in storage systems face challenges in maintaining data reliability when error detection codes are separated and reattached, leading to decreased efficiency and reliability, especially when data is compressed or changed.

Innovation Solution

A storage control apparatus that separates error detection codes from target data, eliminates duplicates, generates new error detection codes for de-duplicated data, and writes these codes back into the memory device, ensuring data integrity and capacity efficiency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If error detection codes are separated from data to enable de-duplication, then capacity utilization is improved, but data reliability deteriorates

Engineering Contradiction:
Improvecapacity utilizationVSAvoiddata reliability
Core Design Contradiction:
Quantity of substanceVSReliability

Solution Approach 1:

The patent segments the data processing into two distinct phases: first separating error detection codes from target data to enable de-duplication, then regenerating error detection codes from the de-duplicated data. This segmentation allows de-duplication to work on pure data while reliability is maintained through code regeneration, resolving the contradiction between capacity utilization and data reliability.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent performs preliminary separation of error detection codes before de-duplication, then performs the action of regenerating new error detection codes after de-duplication. This preliminary and subsequent action sequence ensures that de-duplication operates on clean data without codes, maximizing capacity utilization while the subsequent code regeneration restores data reliability.

Inventive Principle:
Principle #10Preliminary action

2Reliability

If de-duplication is performed on data with error detection codes, then data reliability is maintained, but capacity utilization deteriorates

Engineering Contradiction:
Improvedata reliabilityVSAvoidcapacity utilization
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent extracts error detection codes from the target data before performing de-duplication. By taking out the codes that cause de-duplication inefficiency, the system can perform de-duplication on pure data, achieving high capacity utilization. The extracted codes are then used to verify original data while new codes are generated for de-duplicated data, maintaining reliability.

Inventive Principle:
Principle #2Taking out (Extraction)

3Reliability

If error detection codes are regenerated after de-duplication, then data reliability is improved, but processing complexity increases

Engineering Contradiction:
Improvedata reliabilityVSAvoidprocessing complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system performs self-service by automatically regenerating error detection codes from the de-duplicated data using the same de-duplication apparatus. This self-service mechanism improves data reliability without requiring external intervention or significantly increasing processing complexity, as the code generation uses existing de-duplication infrastructure.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS10084484B2Storage control apparatus and non-transitory computer-readable storage medium storing computer program
Publication Date: 2018.09.25 FUJITSU LTD
  • US10084484B2 patent drawing
  • US10084484B2 patent drawing
  • US10084484B2 patent drawing

AI summary

A storage control apparatus obtains first-code attached data, each having target data to be written and first code information, which includes an error detection code based on the target data and information about a first write destination, attached to the target data. The storage control apparatus then obtains the target data by excluding the first code information from the first-code attached data eliminates duplication of the target data, generates second code information which includes an error detection code for the target data remaining and information about a second write destination, and writes second-code attached data including the second code information into a memory device.