Hyper-scale P2P Deduplicated Storage Using Distributed Ledger
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data deduplication methods are limited to corporate locality, leading to inefficiencies and security concerns when trying to share common data across multiple enterprises, as they often require customers to trust external service providers with their data, which can result in increased costs and vulnerability to data breaches.
Innovation Solution
A hyper-scale, peer-to-peer deduplicated storage system (HSAN) that uses an immutable distributed ledger, such as blockchain technology, to securely share data across multiple organizations by storing only content fingerprints and not metadata, allowing for secure deduplication without compromising user data integrity.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If data is pooled across multiple customers for deduplication, then deduplication efficiency is improved, but data security and trust are worsened
Solution Approach 1:
The patent segments data into discrete chunks and creates individual fingerprints for each chunk. This allows deduplication to operate at the chunk level across multiple customers without exposing complete files or metadata, thereby maintaining deduplication efficiency while enhancing data security through granular isolation.
Solution Approach 2:
The patent introduces an intermediary distributed ledger (blockchain) that mediates between customers and the deduplication system. The ledger stores only fingerprints and location information, acting as a trusted intermediary that enables cross-customer deduplication verification without requiring customers to trust each other or the service provider with their actual data.
2Productivity
If centralized service provider manages deduplication, then deduplication scale is improved, but cost control and customer benefit are worsened
Solution Approach 1:
The patent enables customers to self-verify deduplication through the distributed ledger. Each customer can independently check the ledger to confirm their data chunks are deduplicated across the network, eliminating the need for a centralized service provider to manage and verify deduplication, thereby reducing operational costs and allowing customers to retain more value.
Solution Approach 2:
The distributed ledger serves multiple functions simultaneously: it stores fingerprints for deduplication verification, tracks data location information, provides audit trails for security compliance, and enables cross-customer verification. This multi-functionality reduces the need for separate centralized management systems, lowering overall infrastructure costs.
3Measurement precision
If actual data is shared for deduplication, then deduplication accuracy is improved, but data breach vulnerability is worsened
Solution Approach 1:
The patent extracts only the essential fingerprint information (hash values) and location data from the actual customer data, storing these extracted elements in the distributed ledger while leaving the actual data chunks stored locally with customers. This extraction enables accurate deduplication matching without exposing sensitive data, thereby maintaining deduplication accuracy while eliminating data breach vulnerability.
Data Source
AI summary
One example method includes receiving from a node, in an HSAN that includes multiple nodes, an ADD_DATA request to add an entry to a distributed ledger of the HSAN, the request comprising a user ID that identifies the node, a hash of a data segment, and a storage location of the data segment at the node, performing a challenge-and-response process with the node to verify that the node has a copy of the data that was the subject of the entry, making a determination that a replication factor X has not been met, and adding the entry to the distributed ledger upon successful conclusion of the challenge-and-response process.


