Column-Oriented Data Encoding Using Pattern-Ordered RLE

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In-memory database systems face challenges in efficiently storing and accessing large amounts of data due to limited memory, requiring data compression and CPU-intensive scanning operations, especially in column-oriented storage where frequent column values are too large for caching in a decompressed manner.

Innovation Solution

A method using data mining algorithms to find frequent column patterns, ordering values into a prefix tree, and applying run-length encoding to optimize data compression, while recursively processing precompressed data tuples to maximize compression rates and reduce memory overhead.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If data is compressed to store more data in limited memory, then storage capacity is improved, but access speed and processing efficiency deteriorate

Engineering Contradiction:
Improvestorage capacityVSAvoidaccess speed
Core Design Contradiction:
Quantity of substanceVSSpeed

Solution Approach 1:

The patent segments data into columnar structures and applies different compression techniques to different columns based on their characteristics. Frequently accessed columns are kept in less compressed forms or in memory caches, while less frequently accessed columns use higher compression ratios. This segmentation allows the system to optimize storage capacity for bulk data while maintaining access speed for frequently queried columns.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements local quality by applying varying compression strategies to different parts of the data structure. Hot data (frequently accessed) is stored with lower compression or in RAM, while cold data (infrequently accessed) uses aggressive compression. The system dynamically adjusts compression levels based on access patterns, ensuring that local access requirements are met without compromising overall storage capacity.

Inventive Principle:
Principle #3Local quality

2Quantity of substance

If data is stored in column-oriented manner for compression efficiency, then storage density is improved, but CPU intensive scanning operations worsen performance

Engineering Contradiction:
Improvestorage densityVSAvoidCPU energy consumption
Core Design Contradiction:
Quantity of substanceVSUse of energy by moving object

Solution Approach 1:

The patent applies preliminary action by pre-compressing data during the data loading and ETL (Extract, Transform, Load) phases rather than during query execution. Compression is performed when CPU resources are more readily available and can be dedicated to this task without impacting query performance. The compressed data is then stored and can be efficiently scanned with minimal CPU intervention during actual query operations.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent replaces mechanical CPU-intensive scanning with more efficient access patterns. By organizing data in columnar format with intelligent compression, the system reduces the amount of data that needs to be read and processed. Vectorized query execution and use of SIMD (Single Instruction, Multiple Data) instructions substitute for traditional row-by-row processing, reducing overall CPU energy consumption while maintaining storage density.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Speed

If frequent column values are cached in decompressed manner for fast access, then access speed is improved, but memory consumption increases

Engineering Contradiction:
Improveaccess speedVSAvoidmemory consumption
Core Design Contradiction:
SpeedVSQuantity of substance

Solution Approach 1:

The patent implements dynamic memory management for cached data. The system dynamically adjusts the amount of memory allocated for caching frequently accessed column values based on available memory resources and observed access patterns. When memory pressure is detected, the system evicts less frequently accessed cached values and recompresses them. This dynamic approach allows the system to maintain fast access speeds for hot data while adapting memory consumption to available resources.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent changes parameters of cached data storage by using different compression levels for cached versus non-cached data. Frequently accessed values in the cache may use lighter compression or no compression for maximum access speed, while values that fall out of the cache are recompressed to higher levels. The system monitors access patterns and adjusts compression parameters dynamically, allowing fast access for active cached data while managing overall memory consumption through parameter changes in compression intensity.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS9325344B2Encoding data stored in a column-oriented manner
Publication Date: 2016.04.26 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US9325344B2 patent drawing
  • US9325344B2 patent drawing
  • US9325344B2 patent drawing

AI summary

Data stored in a column-oriented manner is encoded using a data mining algorithm for finding column patterns among a set of data tuples, where each data tuple contains a set of columns, and the data mining algorithm treats all columns and all column combinations and column ordering similarly or in the same manner when looking for column patterns. Column values are ordered occurring in the column patterns based on their frequencies into a prefix tree, where the prefix tree defines a pattern order. The data tuples are sorted according to the pattern order, resulting in sorted data tuples, and columns of the sorted data tuples are encoded using run-length encoding.