Unified Storage System with Multi-threaded Transaction Log
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current High Availability (HA) storage systems require separate solutions for primary data storage, data protection, and analytics, which can lead to performance impacts and increased complexity.
Innovation Solution
A unified system that integrates primary data storage, data protection, and analytics using multi-threaded log writes, nested virtual machine directories, file system cloning, and write gathering across nodes, enabling inline data analytics and efficient data restoration without impacting primary storage performance.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If separate solutions are used for primary data storage, data protection, and analytics, then each function can be independently optimized, but system complexity increases and performance is impacted
Solution Approach 1:
The patent combines primary data storage, data protection, and analytics into a single unified storage system. The storage node simultaneously performs primary storage operations, mirrors data for protection, and executes analytics on the same physical infrastructure, eliminating the need for separate dedicated systems for each function while maintaining all required capabilities
Solution Approach 2:
The storage node is designed to perform multiple functions concurrently: it acts as a primary storage device for data operations, a protection device for mirroring and data safety, and an analytics engine for processing and generating insights. This multi-functional design allows a single system to replace what would traditionally require multiple separate specialized systems
2Reliability
If data is mirrored and tracked for protection to a recovery pool, then data protection is enhanced, but storage performance is impacted
Solution Approach 1:
The system performs data mirroring and protection operations continuously in the background without interrupting primary storage performance. Analytics are also executed continuously on protected data, and the system maintains continuous synchronization between primary and recovery pools through efficient background processes that do not block primary operations
Solution Approach 2:
The patent introduces a buffer cache and asynchronous communication mechanisms as intermediaries between the primary storage operations and the protection/analytics processes. These intermediaries allow primary write operations to complete quickly while protection and analytics proceed in the background, decoupling the performance-critical primary operations from the protective but potentially slower mirroring and analysis processes
3Loss of information
If analytics are gathered on protected data, then data intelligence is improved, but processing time is increased
Solution Approach 1:
The system performs analytics on protected data in the background as data is being mirrored and protected, rather than waiting until protection is complete. This allows analytics to begin processing data as soon as it arrives at the recovery pool, overlapping the protection and analytics processes to eliminate sequential waiting time
Solution Approach 2:
Analytics are executed continuously on protected data as it becomes available in the recovery pool, maintaining an ongoing process of data intelligence generation without interruption or batch processing delays. The system continuously analyzes protected data streams, generating insights in real-time rather than through periodic or on-demand operations
Data Source
AI summary
A unified system provides primary storage and in-line analytics-based data protection. Additional data intelligence and analytics gathered on protected data and prior analytics are stored in discovery points. The disclosed system implements multi-threaded log writes across primary and restore nodes with write gathering across file systems; nested directories such as may be used for storing virtual machine files, where every subdirectory has an associated file system for snapshot purposes; and cloning objects on demand with background metadata and data migration.


