Anonymized Data Codebooks for Compaction and Secure Transmission

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The rapid growth of data storage demand outpaces the capacity to store it, leading to storage and transmission bottlenecks, and existing data compression and encryption technologies are inadequate for securing and efficiently managing anonymized data.

Innovation Solution

A system and method for data compaction and encryption of anonymized data records involve preprocessing datasets into sourceblocks, creating a tally record, and using a distributed data deconstruction engine and library manager to generate a codebook with unique codewords for efficient compaction and encryption, leveraging multiple codebooks and distributed architectures for scalability and security.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If data compression is used to reduce storage demand, then storage capacity is improved, but compression ratios are insufficient for multi-media data and result in data degradation

Engineering Contradiction:
Improvestorage capacityVSAvoiddata quality
Core Design Contradiction:
Quantity of substanceVSManufacturing precision

Solution Approach 1:

The patent segments data into discrete tokens and uses tally records to count token occurrences. By dividing data into manageable token units and using distributed deconstruction engines, the system achieves high compression ratios without degrading data quality, as each token can be precisely reconstructed during decompression.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent creates compact representations (tally records and codebooks) that capture the essential structure and frequency distribution of the original data. These compact copies enable lossless reconstruction of the original data while occupying minimal storage space, achieving both high compression and data fidelity.

Inventive Principle:
Principle #26Copying

2Quantity of substance

If physical storage capacity is increased to meet demand, then storage capacity is improved, but manufacturing capacity cannot keep up with exponential growth in data storage demand

Engineering Contradiction:
Improvestorage capacityVSAvoidmanufacturing capacity
Core Design Contradiction:
Quantity of substanceVSEase of manufacture

Solution Approach 1:

The patent fundamentally changes the parameter of data representation from storing actual data values to storing compact tally records that encode frequency distributions. This parameter transformation enables exponential compression ratios, allowing petabytes of data to be represented in gigabytes of storage capacity.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent transitions from one-dimensional storage (storing data sequentially) to multi-dimensional storage by creating codebooks that organize tokens by frequency and using distributed deconstruction engines that process data in parallel dimensions, dramatically increasing storage efficiency.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Reliability

If data is anonymized to protect privacy and comply with regulations, then data security is improved, but existing encryption technologies are vulnerable to quantum computing attacks

Engineering Contradiction:
Improvedata securityVSAvoidquantum computing vulnerability
Core Design Contradiction:
ReliabilityVSObject-affected harmful factors

Solution Approach 1:

The patent introduces tally records as an intermediary layer between the original data and its storage representation. These tally records contain only frequency counts of anonymized tokens, not the actual data values, providing inherent security against quantum attacks while maintaining data utility for analysis.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent extracts only the essential statistical properties (token frequencies) from the original data, discarding all identifiable information. This extraction creates a secure anonymized representation that is resistant to quantum decryption while preserving the ability to perform data analysis.

Inventive Principle:
Principle #2Taking out (Extraction)

4Ease of operation

If data is transmitted between large data centers, then data accessibility is improved, but transmission bandwidth becomes a bottleneck

Engineering Contradiction:
Improvedata accessibilityVSAvoidtransmission bandwidth
Core Design Contradiction:
Ease of operationVSQuantity of substance

Solution Approach 1:

The patent changes the parameter of data transmission by converting large datasets into compact tally records before transmission. This transformation reduces transmission bandwidth requirements from petabytes to gigabytes while maintaining data accessibility through efficient distributed query processing.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS12449973B2System and method for data compaction and encryption of anonymized data records
Publication Date: 2025.10.21 ATOMBEAM TECH INC
  • US12449973B2 patent drawing
  • US12449973B2 patent drawing
  • US12449973B2 patent drawing

AI summary

A system and method for data compaction and encryption of anonymized data records. A dataset may be pre-processed by dividing into sourceblocks at reasonable intervals and tallying each sourceblock's frequency, creating a tally record of tokens and count values. This tally record may then be anonymized and transmitted to a data deconstruction engine which combined with a library manager creates a codebook and performs optimization techniques on the codebook. The data deconstruction engine and library manager may be distributed across multiple nodes or devices. The received anonymized tally record may be parsed into individual tokens by identifying the tokens with the highest count value. The tokens may then be sent descending order of count value to the library manger where each token may be assigned a codeword. A half-backed codebook is then created using the tokens and each token's unique codeword, before sending the half-backed codebook to a system user.