Scale-Out Storage Node Failover Using Emergency Processing Modes
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In scale-out storage systems for IaaS/PaaS, existing solutions require additional storage resources and increased costs to maintain service level agreements (SLAs) during failures, as they allocate resources only within a node, leading to reduced service quality and increased load on remaining nodes.
Innovation Solution
A storage system with a normal mode and an emergency mode, where the process mode is switched to emergency mode in a second storage node upon failure of a first node, suppressing certain functions like compression and deduplication to reduce processing loads and maintain SLA without additional resources.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If resources are allocated only within a node upon failure, then resource allocation is simple, but the load on remaining nodes increases and service quality is reduced
Solution Approach 1:
The patent extends resource allocation from a single-node dimension to a multi-node cluster dimension. When a storage node fails, the system can allocate resources across multiple remaining nodes in the cluster, not just within the failed node's local resources. This dimensional expansion allows load distribution across the cluster, preventing any single node from becoming overloaded and maintaining service quality without requiring complex intra-node allocation mechanisms.
2Reliability
If extra storage resources are included in each node to maintain performance upon failure, then storage performance SLA is maintained, but the cost is high
Solution Approach 1:
The patent implements universal resource pools at the cluster level that can serve multiple purposes and multiple nodes. Instead of each node requiring dedicated extra resources for failure scenarios, the cluster maintains shared resource pools that can be dynamically allocated to any node experiencing failure. This multi-functional resource allocation maintains storage performance SLAs during failures while reducing the total quantity of storage resources needed compared to per-node redundancy approaches.
3Reliability
If resources are allocated from another logical partition, then storage performance is maintained, but many extra storage resources are needed and cost is high
Solution Approach 1:
The patent segments the storage system into independent storage nodes within a cluster, each capable of operating autonomously. When a node fails, the system allocates resources from other segments (nodes) in the cluster rather than requiring complex logical partition management within a single node. This segmentation simplifies the allocation mechanism by treating each node as a discrete resource unit that can be independently managed and reallocated, reducing overall system complexity while maintaining performance.
Data Source
AI summary
While an extra storage resource required for an operation of IaaS/PaaS is reduced, an SLA on storage performance is maintained even upon a failure. In a storage system including a plurality of storage nodes for providing storage regions for storing data of a computer on which an application is executed, a normal mode to be set in a normal state and an emergency mode in which a predetermined function is suppressed compared with the normal mode are prepared as a process mode for a request for input and output of data. In the storage system, in response to the occurrence of a failure in a first storage node, the process mode is switched to the emergency mode for a second storage node in which the failure does not occur.


