Meta File System Indexing for Big Data Management

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional techniques fail to efficiently manage and analyze Big Data due to its sheer size, volume, and complex structure, leading to difficulties in data storage, retrieval, and analysis, particularly for unstructured data, and lack contextual and semantic information handling.

Innovation Solution

A Meta File System that models Big Data using meta-information representations, enabling efficient data management and analysis by decoupling content from file system structure, using Flash storage for quick metadata access, and creating metadata maps to reduce data transfer needs, allowing for semantically and contextually aware data organization and analysis.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If conventional data storage and management techniques are used, then data can be stored, but the sheer size and complex structure of Big Data makes it difficult to manage and analyze efficiently

Engineering Contradiction:
Improvedata management efficiencyVSAvoiddata structure complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent segments Big Data into structured and unstructured components, applying different management techniques to each. Structured data is organized in traditional databases while unstructured data is managed through content-addressable storage and metadata indexing, enabling efficient processing of complex data structures.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces metadata as an intermediary layer between raw data and user queries. This metadata layer includes contextual information, data types, and relationships, which mediates the complexity of Big Data structures and enables efficient indexing and retrieval without requiring users to navigate complex data structures directly.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Speed

If sequential data access is used, then data can be retrieved, but it is time-consuming and inefficient for large datasets

Engineering Contradiction:
Improvedata retrieval speedVSAvoidaccess time
Core Design Contradiction:
SpeedVSLoss of time

Solution Approach 1:

The patent performs preliminary indexing and categorization of data into structured and unstructured segments before retrieval operations. Metadata is pre-computed and stored with data, enabling direct access to specific data types without sequential scanning, significantly reducing access time for large datasets.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent changes the access parameter from sequential position-based access to content-based access using hashing and metadata indexing. This allows random access to any data element based on its content characteristics rather than its position in a sequence, dramatically improving retrieval speed.

Inventive Principle:
Principle #35Parameter changes

3Adaptability or versatility

If data is stored without contextual information, then storage is simplified, but analysis lacks semantic understanding and cross-correlation capability

Engineering Contradiction:
Improvedata analysis capabilityVSAvoiddata organization structure
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent creates a universal metadata framework that serves multiple functions: data categorization, indexing, contextual enrichment, and cross-correlation enablement. This single metadata layer supports diverse data types and analysis requirements without requiring separate organizational structures for each function.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent embeds multiple levels of contextual information within a nested metadata structure. Data elements contain metadata about their structure, which in turn contains metadata about relationships, and so on. This nested organization allows rich contextual information to be stored efficiently without overwhelming complexity at any single level.

Inventive Principle:
Principle #7Nested doll (Nesting)

4Productivity

If metadata maps are created to reduce data transfer, then data access efficiency improves, but metadata management complexity increases

Engineering Contradiction:
Improvedata access efficiencyVSAvoidmetadata management complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent extracts essential metadata characteristics from large datasets and stores only these summarized attributes in the metadata map. By taking out only the necessary indexing information rather than copying entire data structures, the system reduces metadata management complexity while maintaining efficient data access through the condensed metadata representations.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS9275059B1Genome big data indexing
Publication Date: 2016.03.01 EMC IP HLDG CO LLC
  • US9275059B1 patent drawing
  • US9275059B1 patent drawing
  • US9275059B1 patent drawing

AI summary

A computer implemented method, computer program product, and apparatus for modeling a Big Data dataset, the method comprising creating non-specific representations of the Big Data dataset by representing, as objects in a computer model, non-specific representations including metaInformation, DataSet, BigData and Properties representations and creating non-specific representations of indices, wherein the indices are mapped to one or more key-value pairs.