Thin Database Indexing via Permutation Cyclic Lookups

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Traditional database management techniques are impractical for handling large and complex Big Data sets, as they incur high storage costs and slow query times due to the need for extensive indexing and disk access, which becomes prohibitive in Big Data applications.

Innovation Solution

The implementation of thin database indexing techniques, which encode data in a permutation cyclic structure, allowing for efficient retrieval of row numbers using minimal indexing overhead, reducing disk access and storage requirements by using a permutation function to determine row numbers through cyclic look-ups and shortcuts.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional database indexing techniques are used to handle Big Data sets, then data retrieval capability is maintained, but storage costs and disk access requirements increase significantly

Engineering Contradiction:
Improvedata retrieval capabilityVSAvoidstorage space
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent extracts only the essential indexing information needed for data retrieval, storing minimal index data that points to actual data locations rather than maintaining complete traditional indexes. This extraction approach maintains query capability while dramatically reducing storage requirements for Big Data sets.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

Instead of traditional indexing where index structures store comprehensive data relationships, this patent inverts the approach by using compact index markers that point to data locations, reducing index storage while maintaining retrieval efficiency through reverse lookup mechanisms.

Inventive Principle:
Principle #13The other way round (Inversion)

2Reliability

If traditional database indexing techniques are used to handle Big Data sets, then data retrieval capability is maintained, but disk access time increases

Engineering Contradiction:
Improvedata retrieval capabilityVSAvoiddisk access time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent segments the index into multiple compact parts distributed across the data structure, allowing parallel access and reducing the time needed to locate data. This segmentation enables efficient handling of Big Data without requiring extensive sequential disk access.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent performs preliminary organization of data and index markers during data ingestion, pre-positioning index information in optimized locations. This preliminary action reduces the need for extensive disk seeking and access during query operations, significantly improving retrieval speed.

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If extensive indexing is implemented for Big Data, then query accuracy is maintained, but device complexity increases

Engineering Contradiction:
Improvequery accuracyVSAvoidindexing structure complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent changes the parameters of the indexing structure from traditional complex B-trees or hash indexes to simplified marker-based location systems. This parameter change maintains query accuracy by preserving precise data location information while dramatically reducing structural complexity for Big Data handling.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentEP3036663B1Thin database indexing
Publication Date: 2017.10.04 PIVOTAL SOFTWARE INC
  • EP3036663B1 patent drawingFigure 1
  • EP3036663B1 patent drawingFigure 2
  • EP3036663B1 patent drawingFigure 3

AI summary

Methods, systems, and apparatus, including computer programs encoded on computer storage media, for indexing column values of a database table. One of the methods includes receiving a database query requesting one or more rows of a database table having N rows, each row having an encoded column value, wherein the encoded column values are between 1 and N inclusive, wherein each distinct encoded column value occurs in the column exactly once, wherein each encoded column value represents a corresponding original column value, and wherein the query specifies a first original column value. The first original column value is mapped to a first row number identifying a first row. A row is identified as satisfying the query by chaining through the encoded column values in the table from the first row to identify subsequent rows until a row having an encoded column value matching the first row number is reached.