Thin Database Indexing via Permutation Cyclic Lookups
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Traditional database management techniques are impractical for handling large and complex Big Data sets, as they incur high storage costs and slow query times due to the need for extensive indexing and disk access, which becomes prohibitive in Big Data applications.
Innovation Solution
The implementation of thin database indexing techniques, which encode data in a permutation cyclic structure, allowing for efficient retrieval of row numbers using minimal indexing overhead, reducing disk access and storage requirements by using a permutation function to determine row numbers through cyclic look-ups and shortcuts.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional database indexing techniques are used to handle Big Data sets, then data retrieval capability is maintained, but storage costs and disk access requirements increase significantly
Solution Approach 1:
The patent extracts only the essential indexing information needed for data retrieval, storing minimal index data that points to actual data locations rather than maintaining complete traditional indexes. This extraction approach maintains query capability while dramatically reducing storage requirements for Big Data sets.
Solution Approach 2:
Instead of traditional indexing where index structures store comprehensive data relationships, this patent inverts the approach by using compact index markers that point to data locations, reducing index storage while maintaining retrieval efficiency through reverse lookup mechanisms.
2Reliability
If traditional database indexing techniques are used to handle Big Data sets, then data retrieval capability is maintained, but disk access time increases
Solution Approach 1:
The patent segments the index into multiple compact parts distributed across the data structure, allowing parallel access and reducing the time needed to locate data. This segmentation enables efficient handling of Big Data without requiring extensive sequential disk access.
Solution Approach 2:
The patent performs preliminary organization of data and index markers during data ingestion, pre-positioning index information in optimized locations. This preliminary action reduces the need for extensive disk seeking and access during query operations, significantly improving retrieval speed.
3Measurement precision
If extensive indexing is implemented for Big Data, then query accuracy is maintained, but device complexity increases
Solution Approach 1:
The patent changes the parameters of the indexing structure from traditional complex B-trees or hash indexes to simplified marker-based location systems. This parameter change maintains query accuracy by preserving precise data location information while dramatically reducing structural complexity for Big Data handling.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
Methods, systems, and apparatus, including computer programs encoded on computer storage media, for indexing column values of a database table. One of the methods includes receiving a database query requesting one or more rows of a database table having N rows, each row having an encoded column value, wherein the encoded column values are between 1 and N inclusive, wherein each distinct encoded column value occurs in the column exactly once, wherein each encoded column value represents a corresponding original column value, and wherein the query specifies a first original column value. The first original column value is mapped to a first row number identifying a first row. A row is identified as satisfying the query by chaining through the encoded column values in the table from the first row to identify subsequent rows until a row having an encoded column value matching the first row number is reached.