Object File System for Hybrid Storage Deduplication
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Hybrid storage systems combining primary storage and cloud computing face compatibility issues, leading to loss of robust functionalities like data replication and deduplication when integrating with cloud storage, resulting in cost inefficiencies and scalability limitations.
Innovation Solution
An object file system is implemented to store and manage data within an object store, utilizing a snapshot file system tree structure for efficient data representation, deduplication, and metadata management, allowing for on-demand access and scalability while maintaining compatibility with both primary and cloud storage systems.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If primary storage systems are used to provide robust functionalities like data replication and deduplication, then data management capability is improved, but cost increases and scalability decreases
Solution Approach 1:
The system segments storage functionality into two parts: primary storage systems provide robust data management features (replication, deduplication), while cloud storage provides scalable capacity. The hybrid architecture separates concerns between reliability and scalability, allowing each component to excel at its designated function without forcing the entire system to bear the burden of both requirements simultaneously.
Solution Approach 2:
The system creates a universal hybrid storage architecture that can perform both primary storage functions (data protection, management) and cloud storage functions (scalable capacity, cost-effective storage) through a unified interface. This multi-functional approach allows the system to leverage benefits from both storage types without requiring separate independent systems.
2Device complexity
If cloud computing storage is used to achieve cost savings and scalability, then storage efficiency is improved, but robust functionalities like data replication and deduplication are lost
Solution Approach 1:
The system introduces an intermediary layer that sits between the cloud storage and the user applications. This intermediary maintains and manages the robust data management functionalities (replication, deduplication) that would otherwise be lost when using pure cloud storage. The intermediary acts as a mediator that preserves primary storage capabilities while leveraging cloud storage benefits.
Solution Approach 2:
The system performs data management operations (deduplication, replication) as preliminary actions before data is transferred to or from cloud storage. By preparing and optimizing data beforehand, the system ensures that robust functionalities are maintained even though the actual storage resides in the cloud, preventing functionality loss rather than attempting to restore it afterward.
3Device complexity
If a hybrid storage system combines primary storage and cloud storage, then cost efficiency and scalability are improved, but compatibility issues arise causing functionality loss
Solution Approach 1:
The system changes the interface parameters and communication protocols between primary storage and cloud storage components to ensure compatibility. By standardizing data formats, access methods, and management interfaces, the system enables seamless integration between heterogeneous storage types, allowing them to work together as a unified hybrid system rather than incompatible separate systems.
Data Source
AI summary
Techniques are provided for on-demand creation and/or utilization of containers and/or serverless threads for hosting data connector components. The data connector components can be used to perform integrity checking, anomaly detection, and file system metadata analysis associated with objects stored within an object store. The data connector components may be configured to execute machine learning functionality to perform operations and tasks. The data connector components can perform full scans or incremental scans. The data connector components may be stateless, and thus may be offlined, upgraded, onlined, and/or have tasks transferred between data connector components. Results of operations performed by the data connector components upon base objects may be stored within sibling objects.


