Data Edge Format for Object Storage Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Object storage systems face challenges in efficiently managing and analyzing disjoined, disparate, and malformed data, requiring manual inspection and transformation processes that are time-consuming and costly, hindering the effectiveness of data lakes.

Innovation Solution

The introduction of a data format called 'data edge' that universally represents various data sources, allowing for virtual transformation and aggregation without significant computation, enabling data discovery, organization, compression, and analysis within object storage, while supporting relational queries and text searches without increasing storage footprint.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If data is stored in object storage using traditional formats, then data can be stored with simplicity and scalability, but data becomes disjoined, disparate, and malformed requiring manual inspection and transformation

Engineering Contradiction:
Improvedata storage efficiencyVSAvoiddata analysis capability
Core Design Contradiction:
ProductivityVSEase of operation

Solution Approach 1:

The patent segments data into symbols and their locations, storing them separately. This segmentation allows the data to maintain its structured format while enabling efficient compression and analysis operations without requiring manual transformation.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary data format that sits between traditional object storage and analysis systems. This format acts as a mediator that preserves data simplicity while enabling advanced analysis capabilities through symbolic representation.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If manual inspection and transformation methods are used for data analysis, then data can be processed, but the process becomes time-consuming and costly

Engineering Contradiction:
Improvedata analysis accuracyVSAvoiddata processing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent performs preliminary organization of data into symbols and locations during the storage phase. This preliminary action eliminates the need for time-consuming manual inspection and transformation later, as the data is already structured for efficient analysis.

Inventive Principle:
Principle #10Preliminary action

3Quantity of substance

If data is compressed using traditional methods, then storage footprint may be reduced, but compression ratios are limited and cannot achieve theoretical minimums

Engineering Contradiction:
Improvestorage footprintVSAvoidcompression efficiency
Core Design Contradiction:
Quantity of substanceVSProductivity

Solution Approach 1:

The patent transitions from compressing data in its traditional form to compressing symbolic representations of data. This dimensional change allows compression algorithms to work on the structure and patterns of symbols rather than raw data, achieving superior compression ratios that approach theoretical minimums.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

4Ease of operation

If data is organized into like subsets for efficient use, then data analysis can be improved, but the organization process adds complexity

Engineering Contradiction:
Improvedata organization capabilityVSAvoiddata structure complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent organizes data by grouping similar symbols together, creating homogeneous subsets. This organization method improves data analysis efficiency while maintaining relative simplicity, as symbols of the same type are naturally grouped without requiring complex classification schemes.

Inventive Principle:
Principle #33Homogeneity

Data Source

PatentUS10846285B2Materialization for data edge platform
Publication Date: 2020.11.24 CHAOSSEARCH INC
  • US10846285B2 patent drawing
  • US10846285B2 patent drawing
  • US10846285B2 patent drawing

AI summary

Disclosed are system and methods for processing and storing data files, using a data edge file format. The data edge file separates information about what symbols are in a data file and information about the corresponding location of those symbols in the data file. An index for the data files can be generated according to the data edge file format. Using the data edge index, a materialized view of a result set can be generated in response to a search query for the source data objects stored in object storage.