Data Edge Format for Object Storage Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Object storage systems face challenges in efficiently storing and analyzing disjoined, disparate, and malformed data due to their schema-less nature, requiring manual inspection and transformation processes that are time-consuming and costly.

Innovation Solution

The introduction of a data format called 'data edge' that universally represents various data sources, allowing for virtual transformation and aggregation without significant computation, enabling data discovery, organization, compression, and analysis within the storage layer, with built-in support for relational queries and text searches, and the ability to compress and catalog data to theoretical minimums.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If manual inspection and transformation processes are used for data analysis in object storage, then data can be analyzed, but the process becomes time-consuming and costly

Engineering Contradiction:
Improvedata analysis efficiencyVSAvoidtime for manual inspection
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent implements self-service through automated data discovery, organization, and analysis capabilities embedded directly in the object storage system. The system automatically detects data types, applies appropriate transformations, and generates analytics without human intervention, replacing manual inspection processes with autonomous computational workflows that operate continuously on stored data

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent introduces an intermediary data processing layer between object storage and analysis tools. This intermediary automatically transforms raw data into structured formats, applies metadata tagging, and prepares data for analysis, thereby eliminating the need for manual data preparation and enabling direct analysis of stored data

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If data is stored in a schema-less manner in object storage, then scalability and simplicity are improved, but data becomes disjoined, disparate, and malformed making analysis difficult

Engineering Contradiction:
Improvestorage flexibilityVSAvoiddata organization complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent applies dynamics by implementing adaptive data organization that automatically adjusts to different data types and structures. The system dynamically creates appropriate schemas based on detected data patterns, transforms malformed data into structured formats, and reorganizes disparate data into coherent relationships, thereby maintaining storage flexibility while reducing organizational complexity

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent utilizes parameter changes by automatically detecting data characteristics and transforming data parameters accordingly. The system changes data format parameters, structure parameters, and organization parameters based on the detected data type, converting unstructured data into structured formats that enable efficient analysis while preserving the underlying information

Inventive Principle:
Principle #35Parameter changes

3Productivity

If data is copied out for processing and analysis, then analysis can be performed, but the cycle of Extracting, Transforming, and Loading increases storage requirements and processing overhead

Engineering Contradiction:
Improvedata processing capabilityVSAvoidstorage footprint
Core Design Contradiction:
ProductivityVSQuantity of substance

Solution Approach 1:

The patent applies segmentation by separating data processing functions from data storage locations. The system segments processing operations into discrete transformation steps that can be applied directly to data in object storage, eliminating the need to copy entire datasets for analysis and reducing storage requirements by processing data in-place or with minimal replication

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS11762876B2Data normalization using data edge platform
Publication Date: 2023.09.19 CHAOSSEARCH INC
  • US11762876B2 patent drawing
  • US11762876B2 patent drawing
  • US11762876B2 patent drawing

AI summary

Disclosed are system and methods for processing and storing data files, using a data edge file format. The data edge file format separates information about what symbols are in a data file and information about the corresponding location of those symbols in the data file. Examples convert a source file comprising symbols into a data edge index having a manifest portion, a symbol portion, and a locality portion. The symbol portion contains a sorted unique set of symbols from the source file, and the locality portion contains a plurality of location values referencing the symbol portion. Examples include normalizing structured data from the source file by modifying the locality manifest portion of the data edge file to include a description of at least one nonexistent column empty locality value at a respective position within the locality file representing an omission of data at an associated position in the source file.