Frequency-Based Codebooks for Anonymized Data Compaction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The rapid growth of data storage demand, driven by social media, cloud data centers, and biotech industries, has outpaced the production of physical storage capacity, leading to a bottleneck in data storage and transmission, with existing solutions like data compression offering limited relief and raising concerns about data security, especially with the emergence of quantum computing and stringent privacy regulations.
Innovation Solution
A system and method for data compaction and encryption of anonymized data records, involving preprocessing datasets into sourceblocks, counting their occurrences, anonymizing the tally records, and using a library manager to create a codebook with optimization techniques, allowing for efficient compaction and encryption by assigning unique codewords to tokens based on their frequency, enabling secure and efficient data storage and transmission.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If data compression is used to reduce storage demand, then storage capacity utilization improves, but data security deteriorates due to limited compression ratios and vulnerability to quantum computing attacks
Solution Approach 1:
The patent creates multiple codebooks that are mathematical transformations of each other, where each codebook is a copy with different properties. This allows the system to use one codebook for compression while maintaining security through the existence of multiple equivalent codebooks, resolving the contradiction between compression efficiency and security.
Solution Approach 2:
The patent combines multiple codebooks into a composite structure where each codebook provides different security properties. The composite codebook system achieves both compression (through efficient encoding) and security (through multiple layers of codebook transformations that resist quantum attacks).
2Quantity of substance
If physical storage capacity is increased to meet demand, then data storage capability improves, but manufacturing feasibility deteriorates due to outpaced production capacity
Solution Approach 1:
The patent changes the parameter of data representation from raw binary to coded forms using multiple codebooks. This parameter change achieves effective storage capacity expansion without physical manufacturing, as the same physical storage can hold more effective data through efficient coding schemes.
Solution Approach 2:
The patent segments data into blocks that are encoded using different codebooks from a set of multiple codebooks. This segmentation allows efficient utilization of existing storage capacity by organizing data across multiple codebook structures, effectively increasing storage capability without new manufacturing.
3Reliability
If data is anonymized to protect privacy, then data security improves, but data utility deteriorates due to loss of personal identifying information
Solution Approach 1:
The patent introduces codebooks as intermediary structures that transform data into coded forms. These codebooks act as mediators that protect privacy by obscuring personal identifying information while maintaining data utility, as the coded data can still be processed and analyzed for machine learning applications.
Solution Approach 2:
The patent changes the parameter of data representation through codebook transformations, converting personally identifiable information into coded forms that maintain statistical and analytical properties. This parameter change preserves data utility for analysis while achieving anonymization for privacy protection.
4Productivity
If transmission bandwidth is increased to handle large datasets, then data transmission capability improves, but infrastructure cost deteriorates due to bandwidth bottlenecks
Solution Approach 1:
The patent changes the parameter of data representation to more compact coded forms using multiple codebooks. This parameter change reduces the quantity of data that needs to be transmitted, thereby improving transmission capability without requiring increased bandwidth infrastructure.
Data Source
AI summary
A system and method for data compaction and encryption of anonymized data records. A dataset may be pre-processed by dividing into a plurality of sourceblocks at all reasonable sourceblock lengths, and then counting how many times each sourceblock occurs in the dataset, resulting in a tally record of tokens and their count value. This tally record may then be anonymized and transmitted to a data deconstruction engine which combined with a library manager creates a codebook and performs optimization techniques on the codebook. The received anonymized tally record may be parsed into individual tokens by identifying the tokens with the highest count value. The tokens may then be sent, in descending order of count value, to the library manger where each token may be assigned a codeword. A half-backed codebook is then created using the tokens and each token's unique codeword, before sending the half-backed codebook to a system user.


