Grouped Codebook Management for Unseen Data Compression

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The rapid growth of data storage demand outstrips the capacity to store it, and existing entropy encoding methods inefficiently handle previously-unseen data, leading to inefficient data compaction and transmission bottlenecks, especially with the rise of multimedia data and quantum computing threats to data security.

Innovation Solution

A system and method for codebook management that combines training datasets with similarity scores to create a combined codebook, using mismatch probability estimates to efficiently handle previously-unseen data through entropy encoding, incorporating a secondary encoding process for unmatched data blocks.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If traditional entropy encoding methods are used, then encoding speed is maintained, but data compaction efficiency deteriorates when handling previously-unseen data

Engineering Contradiction:
Improvedata compaction efficiencyVSAvoidhandling previously-unseen data
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The system performs preliminary actions by creating multiple codebooks from different training datasets before actual encoding occurs. When new data arrives, the system can quickly select or combine appropriate codebooks without needing to learn from scratch, enabling efficient handling of previously-unseen data while maintaining compaction effectiveness

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system changes parameters by dynamically selecting and combining different codebooks based on the characteristics of the data being encoded. By adjusting which codebooks are used and how they are combined, the system adapts to different data types and patterns, improving both compaction efficiency and adaptability to new data

Inventive Principle:
Principle #35Parameter changes

2Quantity of substance

If data compression is applied to increase storage capacity, then storage efficiency improves, but transmission bandwidth requirements worsen due to compression overhead

Engineering Contradiction:
Improvestorage capacity utilizationVSAvoidtransmission bandwidth
Core Design Contradiction:
Quantity of substanceVSUse of energy by moving object

Solution Approach 1:

The system segments data into fixed-size blocks that are independently encoded using codebooks. This segmentation allows for more efficient compression ratios while reducing the overhead associated with traditional compression methods, as each block can be optimally encoded without affecting others, thereby improving storage utilization without proportionally increasing transmission bandwidth requirements

Inventive Principle:
Principle #1Segmentation

3Reliability

If existing encryption technologies are used, then data security is maintained, but security deteriorates in the face of quantum computing threats

Engineering Contradiction:
Improvedata securityVSAvoidresistance to quantum computing threats
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The system introduces an intermediary layer of codebook-based encoding that sits between the data and traditional encryption methods. This intermediary encoding provides an additional layer of obfuscation and complexity that would be difficult for quantum computers to break, while still allowing traditional decryption methods to function, thereby maintaining security against both classical and quantum threats

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS12541299B2System and method for codebook management based on data source grouping
Publication Date: 2026.02.03 ATOMBEAM TECH INC
  • US12541299B2 patent drawing
  • US12541299B2 patent drawing
  • US12541299B2 patent drawing

AI summary

A system and method for codebook management is disclosed. Training datasets are obtained from various data sources. A similarity score is generated for each training dataset with reference to the other training datasets. In response to detecting a similarity score above a predetermined threshold for one or more of the other training datasets, a combined codebook is created based on training datasets that have a similarity score above a predetermined threshold. Based on the similarity score, multiple data sources are combined into a group, and the combined codebook is used for the data sources within the group. A mismatch performance metric can be computed for the combined codebook, and a revised combined codebook can be regenerated in response to the mismatch performance metric being above a predetermined threshold.