Grouped Codebook Management for Mixed-Data Compression
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The rapid growth of data storage demand exceeds the capacity to store it, and existing data compression methods are inefficient for mixed data types, especially multimedia data, leading to storage and transmission bottlenecks, while quantum computing poses a threat to current encryption technologies.
Innovation Solution
A system and method for codebook management using neural networks and clustering analysis to group devices based on data similarities, creating optimized codebooks with mismatch probability estimation for efficient data compaction and enhanced encryption.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If data compression is applied to mixed data types, then storage capacity is increased, but compression ratio decreases substantially for multimedia data
Solution Approach 1:
The patent segments data into different types (multimedia, text, structured, unstructured) and applies specialized compression techniques to each segment. This allows optimized compression for each data type rather than using a single generic compression method, thereby maintaining high compression ratios for multimedia data while still achieving overall storage capacity increases.
Solution Approach 2:
The patent dynamically adjusts compression parameters based on data type characteristics. For multimedia data, it uses parameters optimized for visual/audio redundancy removal, while for text data, it uses parameters optimized for character frequency patterns. This parameter adaptation resolves the contradiction by achieving both high compression ratios and increased storage capacity.
2Quantity of substance
If physical storage capacity is added, then storage demand is met, but manufacturing capacity is insufficient
Solution Approach 1:
The patent merges multiple data sets by identifying and extracting common patterns, thereby creating a consolidated representation that requires less total storage capacity. By combining redundant information across different data sources into shared codebooks and pattern libraries, the system achieves increased effective storage capacity without requiring proportional increases in physical manufacturing capacity.
Solution Approach 2:
Instead of storing complete copies of all data, the patent uses copying of common patterns and codebooks that can be referenced multiple times. This allows the system to meet storage demand for multiple data sets while using minimal additional physical storage capacity, as the same pattern copies serve multiple data reconstruction needs.
3Quantity of substance
If entropy encoding is applied to previously-unseen data, then data compaction is achieved, but encoding efficiency decreases
Solution Approach 1:
The patent performs preliminary analysis of data streams to build codebooks and pattern libraries before actual encoding occurs. By pre-processing data to identify common patterns and create lookup tables in advance, the system achieves efficient encoding of previously-unseen data, as the preliminary structures enable rapid pattern matching and code generation without exhaustive analysis during the encoding phase.
4Ease of operation
If data is transmitted across networks, then data access is enabled, but bandwidth limitations constrain application development
Solution Approach 1:
The patent extracts only the essential pattern information and codebook data needed for data reconstruction, transmitting these compact representations across networks rather than complete data sets. This extraction of core information enables data access functionality while dramatically reducing the bandwidth quantity required, as only compressed pattern descriptors need to be transmitted rather than full data content.
Data Source
AI summary
A system and method for codebook management is disclosed. Training datasets are obtained from various data sources. A similarity score is generated for each training dataset with reference to the other training datasets. In response to detecting a similarity score above a predetermined threshold for one or more of the other training datasets, a combined codebook is created based on training datasets that have a similarity score above a predetermined threshold. Based on the similarity score, multiple data sources are combined into a group, and the combined codebook is used for the data sources within the group. A mismatch performance metric can be computed for the combined codebook, and a revised combined codebook can be regenerated in response to the mismatch performance metric being above a predetermined threshold.


