Distributed Virtual Array Storage Processing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Traditional distributed data storage systems face challenges with high write latency, resource contention, and scalability issues, particularly in hyper-converged architectures, where performance is constrained by busy virtual machines and requires expensive and disruptive upgrades.
Innovation Solution
A Distributed Virtual Array (DVA) system that leverages local flash memory and computational power of each server for storage processing, minimizing inter-server coordination and communication, using non-volatile memory for low-latency data storage and caching, and distributing storage nodes to allow independent scalability and reduced noisy-neighbor effects.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If compute-intensive storage processing functions are performed in storage controllers, then data processing capabilities are improved, but write latency increases and system cost increases
Solution Approach 1:
The patent segments storage processing functions by separating compute-intensive operations (compression, deduplication, encryption) from data path operations. Storage processors handle compute-intensive tasks while storage controllers manage data I/O, allowing parallel execution and reducing write latency by avoiding sequential processing bottlenecks.
Solution Approach 2:
The patent introduces storage processors as intermediary components between storage controllers and persistent storage devices. These processors handle compute-intensive storage processing functions, acting as mediators that offload computational tasks from storage controllers, thereby improving overall data processing capability without increasing write latency.
2Productivity
If storage controllers are made more powerful to handle increasing storage processing loads, then processing capability is improved, but system cost increases
Solution Approach 1:
The patent segments storage processing functions by separating compute-intensive operations (compression, deduplication, encryption) from data path operations. Storage processors handle compute-intensive tasks while storage controllers manage data I/O, allowing parallel execution and reducing write latency by avoiding sequential processing bottlenecks.
Solution Approach 2:
The patent introduces storage processors as intermediary components between storage controllers and persistent storage devices. These processors handle compute-intensive storage processing functions, acting as mediators that offload computational tasks from storage controllers, thereby improving overall data processing capability without increasing write latency.
3Adaptability or versatility
If multiple servers share storage controller resources, then resource utilization is improved, but performance degradation occurs due to contention
Solution Approach 1:
The patent segments storage processing functions by separating compute-intensive operations (compression, deduplication, encryption) from data path operations. Storage processors handle compute-intensive tasks while storage controllers manage data I/O, allowing parallel execution and reducing write latency by avoiding sequential processing bottlenecks.
Solution Approach 2:
The patent introduces storage processors as intermediary components between storage controllers and persistent storage devices. These processors handle compute-intensive storage processing functions, acting as mediators that offload computational tasks from storage controllers, thereby improving overall data processing capability without increasing write latency.
Data Source
AI summary
A data storage system includes a plurality of hosts, each of which includes at least one processor and communicates over a network with a plurality of storage nodes, at least one of which has at least one storage device, at least one storage controller, and at least one non-volatile memory. At least one process within a host issues data storage read/write requests. At least one of the hosts has a cache for caching data stored in at least one of the storage nodes. The host writes data corresponding to a write request to at least one remote non-volatile memory and carries out at least one storage processing function; data in the written-to node may then be made available for subsequent reading by a different one of the hosts. Examples of the storage processing function include compression, ECC computation, deduplicating, garbage collection, write logging, reconstruction, rebalancing, and scrubbing.


