Fault-Responsive Workload Placement Across Host Clusters
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing approaches to SSD degradation in data centers result in high costs due to the decommissioning of entire devices despite functional portions, leading to undesirably high costs and resource wastage.
Innovation Solution
Creating a remedial cluster for hosts with faulty SSDs, allowing only stateless workloads and maintaining a pool of defect-free hosts for replacement, thus extending SSD lifespan and reducing costs.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If entire SSD devices are decommissioned when degradation is detected, then data reliability is maintained, but device cost and resource wastage increase significantly
Solution Approach 1:
The patent segments the SSD device into multiple independent regions (degraded regions and healthy regions). Instead of decommissioning the entire device, only the degraded regions are isolated and marked as unavailable, while the healthy regions continue to serve data storage needs. This segmentation allows partial utilization of the SSD, reducing waste and cost while maintaining data reliability through region-level isolation of faults.
2Reliability
If entire SSD devices are decommissioned when degradation is detected, then data reliability is maintained, but resource utilization decreases due to premature disposal of functional portions
Solution Approach 1:
The patent divides the SSD into operable and non-operable regions based on degradation status. The system continues to allocate and use storage capacity from healthy regions, thereby maintaining high resource utilization. This segmentation approach prevents premature disposal of functional portions while ensuring data reliability through isolation of degraded regions.
Solution Approach 2:
The patent changes the operational parameters of the SSD by dynamically adjusting the available capacity based on region health status. As regions degrade, their capacity is progressively reduced or marked unavailable, while healthy regions continue at full capacity. This parameter adjustment allows the system to adapt to degradation without decommissioning the entire device, maintaining both reliability and resource utilization.
3Reliability
If defective portions of SSD are isolated and sequestered from future writes, then data reliability is improved, but the device lifespan is extended with potential risk of additional block errors
Solution Approach 1:
The patent applies preliminary action by proactively identifying and isolating degraded regions before they cause data corruption or device failure. The system monitors SSD health metrics, detects degradation trends, and preemptively marks regions as unavailable. This preliminary isolation prevents future errors while allowing the device to continue operating with remaining healthy regions, thus extending useful lifespan without compromising reliability.
4Reliability
If SSD regions are monitored for read errors and degraded regions are identified, then data reliability is improved, but system complexity increases due to continuous monitoring and management overhead
Solution Approach 1:
The patent implements self-service by enabling the SSD to autonomously monitor its own health status and identify degraded regions through built-in diagnostics and error logging. The system automatically detects read errors, identifies problematic regions, and manages capacity allocation without requiring complex external monitoring infrastructure. This self-service approach improves reliability while minimizing system complexity by leveraging the SSD's inherent capabilities.
Data Source
AI summary
The present disclosure relates to workload placement responsive to fault. One embodiment includes instructions to remove a first host from a first cluster of a software-defined datacenter (SDDC) responsive to a determination of a fault in a hypervisor of the first host, place the first host into a second cluster of the SDDC, wherein the second cluster is designated to run stateless workloads, and add a second host to the first cluster.


