Meta File System Indexing for Big Data Management
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional techniques fail to efficiently manage and analyze Big Data due to its sheer size, volume, and complex structure, leading to difficulties in data storage, retrieval, and analysis, particularly for unstructured data, and lack contextual and semantic information handling.
Innovation Solution
A Meta File System that models Big Data using meta-information representations, enabling efficient data management and analysis by decoupling content from file system structure, using Flash storage for quick metadata access, and creating metadata maps to reduce data transfer needs, allowing for semantically and contextually aware data organization and analysis.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If conventional data storage and management techniques are used, then data can be stored, but the sheer size and complex structure of Big Data makes it difficult to manage and analyze efficiently
Solution Approach 1:
The patent segments Big Data into structured and unstructured components, applying different management techniques to each. Structured data is organized in traditional databases while unstructured data is managed through content-addressable storage and metadata indexing, enabling efficient processing of complex data structures.
Solution Approach 2:
The patent introduces metadata as an intermediary layer between raw data and user queries. This metadata layer includes contextual information, data types, and relationships, which mediates the complexity of Big Data structures and enables efficient indexing and retrieval without requiring users to navigate complex data structures directly.
2Speed
If sequential data access is used, then data can be retrieved, but it is time-consuming and inefficient for large datasets
Solution Approach 1:
The patent performs preliminary indexing and categorization of data into structured and unstructured segments before retrieval operations. Metadata is pre-computed and stored with data, enabling direct access to specific data types without sequential scanning, significantly reducing access time for large datasets.
Solution Approach 2:
The patent changes the access parameter from sequential position-based access to content-based access using hashing and metadata indexing. This allows random access to any data element based on its content characteristics rather than its position in a sequence, dramatically improving retrieval speed.
3Adaptability or versatility
If data is stored without contextual information, then storage is simplified, but analysis lacks semantic understanding and cross-correlation capability
Solution Approach 1:
The patent creates a universal metadata framework that serves multiple functions: data categorization, indexing, contextual enrichment, and cross-correlation enablement. This single metadata layer supports diverse data types and analysis requirements without requiring separate organizational structures for each function.
Solution Approach 2:
The patent embeds multiple levels of contextual information within a nested metadata structure. Data elements contain metadata about their structure, which in turn contains metadata about relationships, and so on. This nested organization allows rich contextual information to be stored efficiently without overwhelming complexity at any single level.
4Productivity
If metadata maps are created to reduce data transfer, then data access efficiency improves, but metadata management complexity increases
Solution Approach 1:
The patent extracts essential metadata characteristics from large datasets and stores only these summarized attributes in the metadata map. By taking out only the necessary indexing information rather than copying entire data structures, the system reduces metadata management complexity while maintaining efficient data access through the condensed metadata representations.
Data Source
AI summary
A computer implemented method, computer program product, and apparatus for modeling a Big Data dataset, the method comprising creating non-specific representations of the Big Data dataset by representing, as objects in a computer model, non-specific representations including metaInformation, DataSet, BigData and Properties representations and creating non-specific representations of indices, wherein the indices are mapped to one or more key-value pairs.


