Opportunistic RDU Offlining for Faulty Datacenter Units
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Datacenter operators face a challenging decision when a resource distribution unit (RDU) becomes faulty, as it may continue to supply resources to some hosts while others become unavailable, making it unclear whether to take the RDU offline for repair, which can impact tenant availability and datacenter capacity.
Innovation Solution
A decision-making process is implemented within the datacenter fabric to assess factors such as tenant availability, datacenter capacity, and host status to determine whether to take a faulty RDU offline for repair, using modules like availability determining, capacity determining, and host availability assessment, potentially employing decision trees and machine learning algorithms.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If the faulty RDU is taken offline for repair, then the reliability of the datacenter system is improved, but the availability of tenant VMs deteriorates
Solution Approach 1:
The system performs preliminary assessment of host status and datacenter capacity before taking the RDU offline. By evaluating whether hosts can be brought back online and assessing datacenter capacity constraints in advance, the system can make an informed decision about offline timing that minimizes impact on tenant VM availability
Solution Approach 2:
The decision-making process is dynamic rather than static. The system continuously monitors host availability status and datacenter capacity, adjusting the decision to offline the RDU based on current conditions. This allows the system to optimize between repair timing and tenant availability based on real-time system state
2Productivity
If the faulty RDU is kept online, then the availability of tenant VMs is maintained, but the reliability of the datacenter system deteriorates
Solution Approach 1:
The system implements feedback by monitoring the status of hosts parented by the faulty RDU and using this information to guide the repair decision. By assessing whether hosts can be brought back online and how many are currently unavailable, the system receives feedback on the actual impact of keeping the RDU online versus taking it offline
3Reliability
If the faulty RDU is taken offline, then the unavailable hosts can be brought back online, but the datacenter capacity is reduced
Solution Approach 1:
The system performs preliminary assessment of datacenter capacity before deciding to offline the RDU. By evaluating current capacity utilization and predicting the impact of bringing hosts back online, the system can determine whether offlining the RDU would create unacceptable capacity constraints
Solution Approach 2:
The decision process considers changes in system parameters including the number of available hosts, datacenter capacity utilization, and tenant VM distribution. By analyzing how these parameters would change if the RDU were taken offline, the system can make an optimized decision
Data Source
AI summary
Embodiments relate to determining whether to take a resource distribution unit (RDU) of a datacenter offline when the RDU becomes faulty. RDUs in a cloud or datacenter supply a resource such as power, network connectivity, and the like to respective sets of hosts that provide computing resources to tenant units such as virtual machines (VMs). When an RDU becomes faulty some of the hosts that it supplies may continue to function and others may become unavailable for various reasons. This can make a decision of whether to take the RDU offline for repair difficult, since in some situations countervailing requirements of the datacenter may be at odds. To decide whether to take an RDU offline, the potential impact on availability of tenant VMs, unused capacity of the datacenter, a number or ratio of unavailable hosts on the RDU, and other factors may be considered to make a balanced decision.


