Opportunistic RDU Offlining for Faulty Datacenter Units

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Datacenter operators face a challenging decision when a resource distribution unit (RDU) becomes faulty, as it may continue to supply resources to some hosts while others become unavailable, making it unclear whether to take the RDU offline for repair, which can impact tenant availability and datacenter capacity.

Innovation Solution

A decision-making process is implemented within the datacenter fabric to assess factors such as tenant availability, datacenter capacity, and host status to determine whether to take a faulty RDU offline for repair, using modules like availability determining, capacity determining, and host availability assessment, potentially employing decision trees and machine learning algorithms.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If the faulty RDU is taken offline for repair, then the reliability of the datacenter system is improved, but the availability of tenant VMs deteriorates

Engineering Contradiction:
ImproveRDU reliabilityVSAvoidtenant VM availability
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The system performs preliminary assessment of host status and datacenter capacity before taking the RDU offline. By evaluating whether hosts can be brought back online and assessing datacenter capacity constraints in advance, the system can make an informed decision about offline timing that minimizes impact on tenant VM availability

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The decision-making process is dynamic rather than static. The system continuously monitors host availability status and datacenter capacity, adjusting the decision to offline the RDU based on current conditions. This allows the system to optimize between repair timing and tenant availability based on real-time system state

Inventive Principle:
Principle #15Dynamics

2Productivity

If the faulty RDU is kept online, then the availability of tenant VMs is maintained, but the reliability of the datacenter system deteriorates

Engineering Contradiction:
Improvetenant VM availabilityVSAvoiddatacenter system reliability
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The system implements feedback by monitoring the status of hosts parented by the faulty RDU and using this information to guide the repair decision. By assessing whether hosts can be brought back online and how many are currently unavailable, the system receives feedback on the actual impact of keeping the RDU online versus taking it offline

Inventive Principle:
Principle #23Feedback

3Reliability

If the faulty RDU is taken offline, then the unavailable hosts can be brought back online, but the datacenter capacity is reduced

Engineering Contradiction:
Improvehost availabilityVSAvoiddatacenter capacity
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The system performs preliminary assessment of datacenter capacity before deciding to offline the RDU. By evaluating current capacity utilization and predicting the impact of bringing hosts back online, the system can determine whether offlining the RDU would create unacceptable capacity constraints

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The decision process considers changes in system parameters including the number of available hosts, datacenter capacity utilization, and tenant VM distribution. By analyzing how these parameters would change if the RDU were taken offline, the system can make an optimized decision

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS10901824B2Opportunistic offlining for faulty devices in datacenters
Publication Date: 2021.01.26 MICROSOFT TECHNOLOGY LICENSING LLC
  • US10901824B2 patent drawing
  • US10901824B2 patent drawing
  • US10901824B2 patent drawing

AI summary

Embodiments relate to determining whether to take a resource distribution unit (RDU) of a datacenter offline when the RDU becomes faulty. RDUs in a cloud or datacenter supply a resource such as power, network connectivity, and the like to respective sets of hosts that provide computing resources to tenant units such as virtual machines (VMs). When an RDU becomes faulty some of the hosts that it supplies may continue to function and others may become unavailable for various reasons. This can make a decision of whether to take the RDU offline for repair difficult, since in some situations countervailing requirements of the datacenter may be at odds. To decide whether to take an RDU offline, the potential impact on availability of tenant VMs, unused capacity of the datacenter, a number or ratio of unavailable hosts on the RDU, and other factors may be considered to make a balanced decision.