Distributed File System Deduplication via Data Characteristic Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In distributed file systems with multiple deduplication storage devices, existing technologies underutilize deduplication services due to redundant data being spread across devices, leading to inefficient storage capacity utilization and load balancing challenges.

Innovation Solution

A device and method that determine the characteristic of data to be stored, identify the most suitable deduplication storage device based on that characteristic, and store redundant data within the same device for deduplication, utilizing a determination unit, identification unit, and storing unit to enhance deduplication awareness and capacity utilization.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Volume of stationary object

If data is stored in a distributed file system across multiple deduplication storage devices, then storage capacity is increased, but deduplication efficiency deteriorates because redundant data is spread across devices

Engineering Contradiction:
Improvestorage capacityVSAvoiddeduplication efficiency
Core Design Contradiction:
Volume of stationary objectVSProductivity

Solution Approach 1:

The system performs preliminary analysis of data characteristics before storing data in the distributed file system. By determining data characteristics in advance and identifying suitable deduplication storage devices beforehand, the system ensures that redundant data is placed in the same device, enabling effective deduplication while maintaining distributed storage capacity

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces an intermediary layer between the distributed file system and storage devices that acts as a mediator. This intermediary determines data characteristics, identifies appropriate deduplication storage devices, and manages data placement to ensure redundant data is stored together, thereby resolving the conflict between distributed storage and deduplication efficiency

Inventive Principle:
Principle #24Intermediary (Mediator)

2Loss of information

If deduplication is performed at the file system level, then deduplication awareness is improved, but storage bandwidth deteriorates compared to block layer deduplication

Engineering Contradiction:
Improvededuplication awarenessVSAvoidstorage bandwidth
Core Design Contradiction:
Loss of informationVSProductivity

Solution Approach 1:

The system applies different deduplication strategies to different data characteristics. By analyzing local data properties and applying appropriate deduplication methods based on those characteristics, the system achieves both deduplication awareness and maintains high storage bandwidth performance

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS10169363B2Storing data in a distributed file system
Publication Date: 2019.01.01 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US10169363B2 patent drawing
  • US10169363B2 patent drawing
  • US10169363B2 patent drawing

AI summary

A device for storing data in a distributed file system, the distributed file system including a plurality of deduplication storage devices, includes a determination unit configured to determine a characteristic of first data to be stored in the distributed file system; an identification unit configured to identify one of the deduplication storage devices of the distributed file system as deduplication storage device for the first data based on the characteristic of the first data; and a storing unit configured to store the first data in the identified deduplication storage device such that the first data and second data being redundant to the first data are deduplicatable within the identified deduplication storage device.