Dark Deployment Rings for Cloud Upgrade Fault Containment

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing cloud-based platforms face inefficiencies and risks in deploying resource unit upgrades without safety mechanisms, leading to widespread failures and unacceptable downtime due to the unpredictable nature of cloud applications and the high-risk, non-backwards compatible upgrades.

Innovation Solution

A system for small-scale deployment of resource unit upgrades, allowing for efficient testing and validation of upgrade variants by selecting a set of resource units based on criteria such as hardware configuration and collecting telemetry data to analyze performance and detect issues, enabling rapid validation and resolution of problems.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If upgrades are broadly deployed to the entire server farm without testing, then deployment speed is improved, but system reliability deteriorates due to unlimited fault domain exposure

Engineering Contradiction:
Improvedeployment speedVSAvoidsystem reliability
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent segments the server farm into multiple deployment rings with different scopes (e.g., single resource unit, rack, data center, entire farm). Upgrades are deployed incrementally through these rings, allowing testing at each stage before full deployment. This segmentation resolves the contradiction by enabling both rapid deployment (through automated multi-ring progression) and reliability (through controlled fault domain exposure at each ring level).

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements preliminary testing in controlled deployment rings before full-scale deployment. A canary resource unit or small subset is upgraded first to validate the upgrade without exposing the entire system. This preliminary action allows deployment speed to be maintained through automation while ensuring reliability by catching issues before broad deployment.

Inventive Principle:
Principle #10Preliminary action

2Productivity

If unvalidated upgrades are deployed at full scale, then deployment completeness is improved, but fault detection capability worsens due to overwhelming scope

Engineering Contradiction:
Improvedeployment completenessVSAvoidfault detection capability
Core Design Contradiction:
ProductivityVSDifficulty of detecting and measuring

Solution Approach 1:

By dividing the deployment into sequential rings with increasing scope, the patent makes fault detection manageable. Each ring represents a controlled subset where faults can be easily detected and attributed to specific upgrade issues. Telemetry data is collected at each ring level, enabling precise fault detection without the noise of full-scale deployment.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements continuous telemetry data collection and analysis at each deployment ring. This feedback mechanism monitors resource unit health metrics, error rates, and performance indicators specific to each ring's scope. When anomalies are detected, the system can halt progression to the next ring and investigate the specific fault, maintaining both deployment completeness and fault detection capability.

Inventive Principle:
Principle #23Feedback

3Device complexity

If upgrades are deployed without deployment rings, then process simplicity is improved, but recovery complexity worsens due to widespread failures

Engineering Contradiction:
Improveprocess simplicityVSAvoidrecovery complexity
Core Design Contradiction:
Device complexityVSEase of repair

Solution Approach 1:

The deployment ring structure creates natural isolation boundaries that simplify recovery. If an upgrade fails at a specific ring, only the resource units in that ring and previous rings are affected, not the entire server farm. This segmentation maintains process simplicity through automated deployment while dramatically reducing recovery complexity by limiting the scope of potential failures.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent prepares rollback mechanisms and recovery procedures in advance for each deployment ring. Before deploying an upgrade to a new ring, the system ensures that recovery options are available and tested. This beforehand cushioning maintains simple deployment processes while ensuring that recovery operations remain manageable even if failures occur at any ring level.

Inventive Principle:
Principle #11Beforehand cushioning (Prior cushioning)

Data Source

PatentEP4396672B1Dark deployment of infrastructure cloud service components for improved safety
Publication Date: 2025.08.27 MICROSOFT TECHNOLOGY LICENSING LLC
  • EP4396672B1 patent drawingFigure 1
  • EP4396672B1 patent drawingFigure 2A
  • EP4396672B1 patent drawingFigure 2B

AI summary

The techniques disclosed herein enable systems to safely deploy a plurality of upgrade variants to different resource units that provide a service by utilizing small-scale deployment and validation. To deploy upgrade variants, a system receives a selection of upgrade variants from a feature group and automatically selects an appropriate set of resource units at which to deploy the upgrade variants. The system is further configured to collect and analyze telemetry data from the set of resource units to determine if any problems have occurred as a result of the deployed upgrade variants. By analyzing the telemetry data, the system can also identify one or more upgrade variants that are causing the problems. In response, the system can remove the identified variants and proceed with deployment of the remaining upgrade variants.