Data Edge File Format for Object Storage Analytics
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Object storage systems face challenges in efficiently managing and analyzing disjoined, disparate, and malformed data, requiring manual inspection and transformation, which is time-consuming and costly, hindering the effectiveness of data lakes.
Innovation Solution
The implementation of a data format called 'data edging' that separates symbols from their locations within files, enabling efficient compression, organization, and analysis without the need for Extract, Transform, Load (ETL) processes, and supporting both relational queries and text searches, while maintaining a compact storage footprint.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If manual inspection and transformation methods (Hadoop, Redshift, Elastic) are used to analyze data in object storage, then data analysis capability is achieved, but time consumption and cost increase significantly
Solution Approach 1:
The patent applies self-service by enabling object storage to perform its own data analysis operations natively. The storage system includes processing units that can directly execute queries, transformations, and analytics on stored data without requiring external manual inspection tools. This allows the storage system to serve its own analytical needs, eliminating the time-consuming ETL processes and manual analysis previously required.
2Adaptability or versatility
If data is stored in disjoined, disparate, and schema-less manner in object storage, then storage flexibility and scalability are improved, but data organization and analysis become complicated and costly
Solution Approach 1:
The patent applies segmentation by dividing data into discrete objects with structured metadata. Each object contains organized elements including data payloads, descriptors, and hierarchical relationships. This segmentation allows flexible storage of diverse data types while maintaining organized structure through defined object schemas, reducing the complexity of data organization compared to completely schema-less storage.
Solution Approach 2:
The patent applies parameter changes by introducing structured metadata parameters and hierarchical organization parameters to the stored objects. These parameters enable efficient querying, filtering, and analysis without restricting storage flexibility. The system can adjust organizational parameters based on data types and access patterns, balancing flexibility with manageability.
3Productivity
If data is copied out for processing and analysis, then data manipulation capability is achieved, but storage footprint increases and processing efficiency decreases
Solution Approach 1:
The patent applies the intermediary principle by introducing a data lakehouse architecture that acts as a mediator between object storage and processing systems. This intermediary layer provides structured access and transformation capabilities without requiring complete data copying. The system can perform in-place transformations and selective data retrieval, reducing the storage footprint associated with multiple data copies while maintaining manipulation capabilities.
Data Source
AI summary
Disclosed are system and methods for processing and storing data files, using a data edge file format. The data edge file separates information about what symbols are in a data file and information about the corresponding location of those symbols in the data file. The described technique for converting a source file comprising symbols into a data edge file includes: generating a locality file of symbol location from the source file to identify locations of the symbols in the source file, generating a symbol file to identify symbols in the source file, and then modifying the locality file of symbol location to associate each symbol from the symbol file with a location in the source file.


