Reference Codebook Deduplication for Secure Compacted Data
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The rapid growth of data storage demand has outpaced the capacity to store it, leading to storage bottlenecks and bandwidth constraints, especially with the rise of multimedia data and the threat of quantum computing, necessitating a secure deduplication solution for compacted data.
Innovation Solution
A system and method for secure deduplication using a data deconstruction engine, data reconstruction engine, library manager, and reference codebook that compares sourceblocks against a codebook to identify duplicates, creating unique reference codes for non-duplicates and storing them in a codeword storage.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If additional physical storage capacity is added, then storage demand is met, but storage capacity production cannot keep up with exponential data growth demand
Solution Approach 1:
The patent creates a virtual copy of the data library in memory, allowing the system to reference and compare data blocks without physically storing duplicate copies on disk. This virtual copying mechanism enables rapid deduplication identification while minimizing actual physical storage requirements, addressing the contradiction between meeting storage demand and the limitations of physical storage production capacity
2Quantity of substance
If data compression is applied, then storage capacity is doubled, but compression ratios decrease for multimedia data and result in data degradation
Solution Approach 1:
The patent segments data into fixed-size blocks and creates a library of unique blocks with references. This segmentation approach allows exact duplication identification without applying compression algorithms to the data itself, thereby maintaining 100% data integrity while achieving significant storage reduction through elimination of redundant blocks rather than through compression ratios that degrade multimedia quality
Solution Approach 2:
The system creates a virtual copy of the data library in memory for rapid comparison operations. This virtual copying enables the system to identify duplicates without physically compressing or transforming the original data, preserving data integrity while achieving storage efficiency through reference-based deduplication
3Manufacturing precision
If data is stored in uncompressed form, then data integrity is maintained, but transmission bandwidth requirements increase tremendously
Solution Approach 1:
The patent creates a compact virtual representation of the data library in memory that serves as a reference guide for transmission. When data needs to be transmitted, the system can reference this compact library copy to identify and transmit only unique blocks or references to existing blocks, dramatically reducing transmission bandwidth requirements while maintaining data integrity through the reference system
4Reliability
If traditional encryption is used, then data security is provided, but encryption technologies are vulnerable to quantum computing threats
Solution Approach 1:
The patent introduces an intermediary layer - the data library with unique blocks and reference codes - that sits between the original data and storage/transmission channels. This library structure provides inherent security through its reference system, where data can be verified and authenticated through the library references without relying on traditional encryption algorithms that are vulnerable to quantum attacks. The library itself acts as a quantum-resistant security mechanism
Data Source
AI summary
A system and methods for secure deduplication of compacted data comprising a data deconstruction engine, a data reconstruction engine, a library manager, a reference codebook, and a codeword storage which performs simultaneous compaction and deduplication of data sets. A data set may be comprised of one or more sourcepackets which may be optimally deconstructed into a plurality of sourceblocks and wherein each sourceblock may be compared against a reference codebook that contains key-value pairs of a sourceblock and its associated reference code in order to determine if a received sourceblock is a duplicate of data already stored within the reference codebook. Non-duplicate sourceblocks can have a reference code algorithmically created and stored in the reference codebook, thereby ensuring that when a duplicate sourceblock is received, it will not be stored as duplicated data.


