Index File Format for Object Storage Data Compression
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Object storage systems face challenges in efficiently managing and analyzing disjoined, disparate, and malformed data due to their schema-less nature, leading to time-consuming and costly manual processes for data extraction, transformation, and analysis.
Innovation Solution
A data format that universally represents various data sources, including text, images, and videos, by separating symbols from their locations, enabling virtual transformation and aggregation without significant computation, and supporting relational queries and text searches, while reducing storage footprint and allowing for self-descriptive export to other formats.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If data is stored in object storage using traditional schema-less formats, then storage flexibility and scalability are improved, but data analysis efficiency and query capability deteriorate
Solution Approach 1:
The patent segments data into symbols and their locations, creating separate index structures for each. This segmentation enables efficient querying and analysis while preserving the flexibility of storing diverse data types in object storage, resolving the contradiction between storage flexibility and analysis efficiency
Solution Approach 2:
The patent introduces an intermediary indexing layer that sits between the raw object storage data and the analysis queries. This intermediary structure (symbol files and locality files) enables efficient data discovery and analysis without requiring manual ETL processes, while the underlying storage remains flexible and schema-less
2Productivity
If manual inspection or transformation is performed on disjoined and malformed data, then data analysis capability is improved, but time consumption and cost increase
Solution Approach 1:
The patent performs preliminary actions by automatically creating symbol files and locality files during data ingestion, before any analysis is needed. This preliminary indexing eliminates the need for time-consuming manual inspection and transformation later, as the data is already organized and searchable
Solution Approach 2:
The system performs self-service by automatically discovering, organizing, and indexing data without requiring manual intervention. The indexing process autonomously handles malformed and disjoined data, eliminating the need for costly and time-consuming manual transformation processes
3Adaptability or versatility
If data is stored in disjoined and disparate formats, then storage simplicity and scalability are improved, but data organization and discovery capability deteriorate
Solution Approach 1:
The patent creates a universal indexing system that works with all data types stored in object storage, regardless of their original format or structure. The symbol files and locality files provide a unified view of disparate data, enabling efficient discovery and organization while preserving storage scalability and simplicity
4Productivity
If ETL processes are performed for data transformation, then data analysis readiness is improved, but processing complexity and cost increase
Solution Approach 1:
The patent extracts only the essential information needed for analysis (symbols and their locations) during the indexing phase, separating this from the actual data transformation processes. This extraction eliminates the need for complex ETL pipelines, as the indexed data is immediately ready for analysis queries
Data Source
AI summary
Apparatus, methods, and computer-readable media for providing frameworks for data source representation and compression using an index file format are disclosed herein. The index file format separate information about symbols in a data file and information about the corresponding location of those symbols in the data file. The described techniques provide mechanisms for reducing the size associated with the representation of the symbols information and/or the size associated with the representation of the location information.


