Dynamic Failure Domain Configuration in Storage Systems
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current storage systems face challenges in dynamically forming failure domains that can adapt to changing configurations and policies, leading to potential data loss due to inadequate redundancy and resource distribution across blades in a storage system.
Innovation Solution
The solution involves dynamically identifying and configuring failure domains within a storage system using a failure domain formation policy that specifies rules for maximum blades, chassis, network hops, bandwidth, storage capacity, and age, ensuring data redundancy and resilience by distributing data across multiple blades and chassis, and adjusting configurations based on changes in system topology and policy updates.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If failure domains are statically configured in storage systems, then system simplicity is maintained, but the system cannot adapt to changing configurations and policies, leading to potential data loss
Solution Approach 1:
The patent implements dynamic failure domain configuration where the system automatically adjusts failure domain membership based on current system state and policy requirements. The controller continuously monitors blade failures, capacity changes, and policy updates, then dynamically reconfigures failure domains to maintain optimal data protection without requiring manual intervention or fixed static assignments.
Solution Approach 2:
The system employs self-service mechanisms where the storage controller autonomously identifies appropriate failure domain configurations, performs necessary data migration, and maintains redundancy requirements without external input. The system automatically detects topology changes and policy modifications, then self-adjusts the failure domain structure to comply with new requirements while preserving data integrity.
2Reliability
If data is distributed across multiple blades and chassis with dynamic reconfiguration, then data redundancy and resilience are improved, but system complexity and reconfiguration overhead increase
Solution Approach 1:
The system implements continuous feedback loops where the controller monitors system state, failure domain composition, and policy requirements. Based on this feedback, the controller automatically determines when reconfiguration is needed and executes appropriate data migration operations. The feedback mechanism ensures that redundancy requirements are continuously maintained while minimizing unnecessary reconfiguration operations.
Solution Approach 2:
The system performs preliminary actions by pre-planning data migration paths and identifying target locations before actual failures occur. When policy changes or topology modifications are detected, the system proactively reconfigures failure domains and migrates data in advance, ensuring that redundancy requirements are maintained without interrupting ongoing operations or requiring reactive emergency responses.
3Reliability
If failure domain configuration is dynamically adjusted based on policy changes, then compliance with data protection policies is improved, but processing time and operational overhead increase
Solution Approach 1:
The system employs periodic evaluation of policy compliance and failure domain configuration. Rather than continuously monitoring and reacting to every possible change, the system periodically assesses whether current configurations meet policy requirements and only initiates reconfiguration when necessary. This periodic approach reduces processing overhead while ensuring that compliance is maintained through scheduled verification and adjustment.
Data Source
AI summary
Dynamically forming a failure domain in a storage system that includes a plurality of blades, each blade mounted within one of a plurality of chassis, including: identifying, in dependence upon a failure domain formation policy, an available configuration for a failure domain that includes a first blade mounted within a first chassis and a second blade mounted within a second chassis, wherein each chassis is configured to support multiple types of blades; and creating the failure domain in accordance with the available configuration.


